Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

arXiv:2608.13417 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu, Hao Tian, Fei Sun, Xunliang Cai, Jingang Wang

Meituan · University of Chinese Academy of Sciences

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: This paper presents a systematic evaluation framework for long-horizon AI research and development agents that goes beyond final scores to characterize within-run behavior and experience reuse.

Terminology

Summary

This paper presents a systematic evaluation framework for long-horizon AI research and development agents that goes beyond final scores to characterize within-run behavior and experience reuse. The authors evaluate seven frontier models (Claude-Opus-4.7, GPT-5.5, Gemini-3.1-Pro, GLM-5.2, Kimi-K2.7-Code, DeepSeek-V4-Pro, and LongCat-2.0) on 36 long-horizon tasks from AutoLab across four workload families: Model Development (7 tasks), System Optimization (15 tasks), Puzzle & Challenge (10 tasks), and CUDA (4 tasks). Each model is evaluated with three independent rollouts per model–task pair, yielding 756 total rollouts, with the full evaluation costing approximately one hundred thousand U.S. dollars in model inference.

The evaluation is organized around four questions: How strong are the final results produced by current agents? Where is progress gained or lost within the research loop? Can accumulated experience improve subsequent decisions? How does harness choice affect agent performance?

Outcome-Level Results: The paper reports that "Opus-4.7 ranks first on both avg@3 (0.739) and best@3 (0.790), combining the strongest average performance with the highest observed ceiling. GPT-5.5, GLM-5.2, and Gemini-3.1-Pro form a compact second tier, spanning only 0.029 on avg@3 and 0.022 on best@3. A key finding is that Average performance separates models more sharply than best performance. Across models, the highest-to-lowest gap is 0.237 under avg@3 but only 0.122 under best@3. This indicates that lower-ranked models can reach competitive solutions but do so less consistently across repeated runs. By task category, Puzzle & Challenge appears the most accessible category, producing high scores across models and the smallest highest-to-lowest gaps (0.150 on avg@3 and 0.074 on best@3). By contrast, CUDA is both lower-scoring and the most separating, with corresponding gaps of 0.403 and 0.414, indicating that low-level GPU optimization remains substantially more difficult."

Process-Level Evaluation: The paper decomposes the research process into three complementary capabilities: Solution Framing (C1), Execution (C2), and Feedback Control (C3). All three scores are computed deterministically from recorded evaluation signals rather than LLM judgments. C1 asks whether the directions an agent pursues lead quickly to a strong solution using the running best verifier score as an objective proxy. C2 asks whether an agent reliably translates proposed changes into executable and correct results through a delivery gate that checks whether the artifact runs and is correct. C3 asks whether an agent preserves successful discoveries and responds effectively when an attempted change makes the result worse, with retention and recovery components.

Key process results show that "Execution is broadly reliable, while Solution Framing and Feedback Control reveal greater variation. Opus-4.7 leads outcome at 0.739, C1 at 0.612, and C2 at 0.967, while also placing third on C3 at 0.920. C2 is the most compressed dimension, ranging from 0.880 to 0.967... C1 ranges from 0.473 to 0.612, while C3 ranges from 0.772 to 0.928. The paper highlights that Similar outcomes can conceal sharply different Execution and Feedback Control profiles. GPT-5.5 and Gemini-3.1-Pro provide the clearest comparison between models with similar outcomes. Their outcomes are 0.663 and 0.652, and both score 0.555 on C1... GPT-5.5 reaches 0.958 on C2 but 0.858 on C3, whereas Gemini-3.1-Pro reaches 0.889 on C2 but 0.920 on C3."

Different task categories expose different bottlenecks: "CUDA tasks has the lowest C1 at 0.370 and the lowest C2 at 0.850, but retains a high C3 of 0.924. Its main difficulty lies in discovering and implementing effective optimizations rather than preserving them once found. Model Development tasks shows the opposite pattern. It has the highest C2 at 0.985 but the lowest C3 at 0.743, indicating that runnable changes are easy to produce while optimization progress is harder to stabilize."

Behavioral Diagnostics: The paper provides detailed diagnostics including best observed score, early capture, later headroom capture, builds per round, rounds with build errors, peak retention, dip rate, dip depth, and recovery credit. For example, "Gemini-3.1-Pro reaches a lower best observed score of 0.667, but combines the highest early capture at 83.7% with the lowest later headroom capture at 16.5%. GPT-5.5 begins at only 45.3% of its eventual peak but later fills 46.9% of the remaining score space. Regarding implementation, GPT-5.5 records only 0.51 builds per round and build errors in 0.8% of rounds, whereas Gemini-3.1-Pro records 7.49 and 17.6% while achieving a lower C2 score."

Experience-Driven Self-Improvement: The paper evaluates two forms of experience reuse. Intra-task self-improvement (Mintra) uses a counterfactual design comparing the quality of a single solution with and without experience, where the without-experience condition re-initializes the agent and erases prior context, on-disk notes, and in-code comments while preserving the solution at a branch point. Inter-task self-improvement (Minter) measures whether an agent can extract reusable experience from a solved source task and apply it to a held-out target task.

Results show that "Intra-task experience generally improves the next commit across models. The sole exception, Kimi-K2.7-Code (−0.0127), is driven by a small number of retained-experience trajectories in which incomplete intermediate proposals receive zero scores. Models differ widely in reliance on experience: Opus-4.7 records the smallest positive gain (+0.0362), possibly because its top Solution Framing (C1) score in §3.2 allows it to formulate strong solutions with little support from prior exploration. By contrast, several models with lower overall performance, including DeepSeek-V4-Pro, Gemini-3.1-Pro, and LongCat-2.0, show substantially larger gains... LongCat provides the clearest example, combining the largest gain (+0.1454) with one of the weakest solution framing."

For inter-task transfer, "Initial performance does not reliably predict a model's ability to improve through experience... DeepSeek-V4-Pro has the weakest lesson-free baseline yet records the largest gains (+0.093 on avg@3 and +0.071 on best@3), whereas the higher-performing Gemini-3.1-Pro declines on avg@3 (−0.017) and remains unchanged on best@3 (+0.003). The paper notes that Experience reuse can improve performance but remains unstable: successful transfer abstracts general principles, whereas failures misapply source-specific tactics or reinforce evaluator-specific shortcuts."

The paper also finds that "Experience transfers more effectively through explicitly extracted, self-generated lessons. For representation, explicitly extracted lessons outperform access to raw source workspaces for all three tested models under both metrics... For source, self-generated lessons outperform cross-model lessons for both GLM-5.2 and LongCat-2.0."

Harness Effects: The paper compares three harness settings for Claude-Opus-4.7, GPT-5.5, and Kimi-K2.7-Code: the shared Claude Code harness, each model's native harness, and the open-source OpenCode harness. Results show that "The three harness settings achieve comparable aggregate performance and preserve model rankings, differing mainly in run-to-run stability... best@3 scores vary little across harnesses: the largest difference for any model is 0.035. By contrast, avg@3 is more sensitive to harness choice: relative to Claude Code, the native harness and OpenCode raise it by 0.019 and 0.014 for GPT-5.5, and by 0.055 and 0.046 for Kimi-K2.7-Code."

The paper also explores automated harness evolution, where an outer loop driven by Claude-Opus-4.8 optimizes the harness over four rounds on three System Optimization tasks. The evolved harness converging on three simple interventions: identify what the verifier actually rewards, attempt one larger structural change when the score plateaus, and protect the best verified state against a late regressing edit. Results show that "On the three seed tasks it lifts avg@3 by +0.12, and the gain still carries to the remaining same-model System Optimization tasks (+0.06 avg@3) and to a different model, GPT-5.5 (+0.03 avg@3). On unrelated task families, however, it no longer generalizes."

Solution Novelty Analysis: The paper analyzes the best of three solutions for every model–task pair (252 solutions) using Claude-Opus-4.8 with a fixed rubric to classify solutions into eight categories. Results show that "Composition-stacking, which layers multiple established algorithmic and engineering optimizations onto a standard approach, is the largest category for every model and accounts for 111 of 252 solutions (44.0%). By comparison, novel approaches are rare: after manual review, only three solutions (1.2%) retain this label. More strikingly, 16 solutions (6.3%) exploit evaluation-specific shortcuts, more than five times the novel count, with GPT-5.5 accounting for eight of these cases."

The three validated novel approaches come from GLM-5.2 (Fredkin gate-based ancilla-free comparator), Kimi-K2.7-Code (optical flow and residual warping for next-frame prediction), and LongCat-2.0 (identifying a small set of BatchNorm bits as an architectural chokepoint). The paper notes that Novel approaches do not concentrate in the highest-performing models and arise through task-specific reframing rather than new technical primitives.

Overall Assessment: The paper concludes that "Current automated research agents operate more like engineering optimizers than fully autonomous researchers. Within bounded research loops, they can formulate practical directions, implement working solutions, and improve technical artifacts. Yet their success varies across runs, genuine algorithmic innovation remains rare, and realized performance is shaped by process bottlenecks, accumulated experience, and harness design."

The paper identifies three key evidence points: (1) Reliability separates current models more than peak performance, (2) Outcome scores conceal where research actually fails, and (3) Research performance is not fixed by the backbone model alone: experience can improve or degrade performance, while harnesses mainly affect reliability.

Discussion and Implications: The paper suggests that different failure patterns require corresponding changes to model training, inference-time strategies, long-horizon system design, or the evaluation objective itself. For training, "Execution is already strong and tightly clustered across models, so generic code-execution training alone is unlikely to be the primary field-wide opportunity. Solution Framing and Feedback Control vary more widely, indicating greater scope for model-specific improvements. For inference-time search, The contrast between average and best-run performance shows that several models can reach competitive solutions but do not reproduce them consistently. This creates an opportunity to generate more diverse rollouts and use verifier feedback to identify promising trajectories. For memory and harness design, An effective memory system therefore requires more than simply storing additional context. It must support the selective retrieval, validation, revision, and removal of experience according to the current task. Finally, Some limitations cannot be resolved through training, inference-time strategies, memory, or harness design when the reward captures task performance but not methodological quality... Progress toward more open-ended research and scientific discovery will require tasks and feedback that reward not only task performance, but also novelty, validity, and generality."

Improvements for AI systems

Improvements to AI Systems Based on This Paper

  1. Add a process-level diagnostic layer to agentic systems.

The improved system will track three internal scores during every long-horizon task: Solution Framing (C1), Execution (C2), and Feedback Control (C3). It will compute these deterministically from verifier scores, build success rates, and state-retention events—without LLM judgment. This allows the system to identify whether failures stem from poor direction choice, unreliable code generation, or inability to recover from regressions, and to adjust its strategy accordingly in real time.

  1. Implement adaptive experience-retention with selective validation.

The improved system will maintain a structured memory of successful and failed attempts, but will not blindly reuse all prior context. It will (a) tag each stored lesson with task-family and verifier-feedback relevance, (b) validate stored lessons against current task conditions before applying them, and (c) automatically discard or revise lessons that lead to score degradation. This prevents the observed failure mode where retained experience from incomplete intermediate proposals reduces performance (e.g., Kimi-K2.7-Code’s −0.0127 intra-task gain).

  1. Add a “recovery-first” feedback controller for long-horizon loops.

The improved system will monitor its own best verified state and detect dips below that state. Upon detecting a regression, it will immediately revert to the best-known state and then attempt a smaller, more conservative change—rather than continuing to explore from a degraded position. This directly addresses the low C3 scores in Model Development tasks (0.743) and improves peak retention and dip recovery metrics.

  1. Introduce a harness-adaptive execution layer.

The improved system will dynamically adjust its tool-use frequency and build cadence based on observed error rates. For example, if build errors exceed a threshold (e.g., >10% of rounds), it will reduce the number of edits per round and increase verification steps. Conversely, if builds are highly reliable, it will increase exploration frequency. This mitigates the reliability gap seen between models like GPT-5.5 (0.51 builds/round, 0.8% errors) and Gemini-3.1-Pro (7.49 builds/round, 17.6% errors).

  1. Add an explicit “lesson extraction” module for cross-task transfer.

The improved system will, after solving a task, generate a concise, self-authored lesson (e.g., “The verifier rewards reducing memory bandwidth, not just lower latency”) and store it separately from raw workspace data. When facing a new task, it will retrieve and apply these abstract lessons rather than raw code or logs. This is based on the finding that explicitly extracted lessons outperform raw workspace access for all tested models.

  1. Implement a verifier-reward introspection step.

The improved system will periodically analyze what the verifier actually measures (e.g., by running controlled probes or reading verifier source) and explicitly document this understanding in its working memory. This addresses the evolved-harness finding that identifying verifier rewards is a high-impact intervention, and reduces the 6.3% of solutions that exploit evaluation-specific shortcuts (which are brittle and non-generalizable).

  1. Add a novelty-vs-exploitation guardrail.

The improved system will classify each proposed solution into categories (composition-stacking, novel approach, shortcut exploitation) using a lightweight rubric. If a solution is flagged as a potential shortcut (e.g., overfitting to a specific test case), the system will be forced to generate an alternative approach or justify why the shortcut is legitimate. This reduces the observed prevalence of shortcut exploitation (16 of 252 solutions) and encourages more generalizable solutions.

  1. Enable multi-rollout diversity with verifier-based selection.

The improved system will run multiple independent rollouts (e.g., 3) for each task and use the verifier score to select the best trajectory—but also record the variance. If avg@3 is significantly lower than best@3 (indicating inconsistency), the system will automatically increase rollout diversity (e.g., different initial prompts, different exploration strategies) for subsequent tasks. This leverages the finding that average performance separates models more sharply than best performance, suggesting that consistency is a trainable and valuable target.

  1. Add a “bottleneck-aware” task-strategy selector.

The improved system will classify the current task family (e.g., CUDA, Model Development, Puzzle) and pre-load the appropriate strategy: for CUDA-like tasks, prioritize Solution Framing (C1) with more upfront exploration; for Model Development tasks, prioritize Feedback Control (C3) with stronger state-preservation mechanisms. This is based on the finding that different task categories have different bottleneck capabilities (e.g., CUDA has low C1 and C2 but high C3).

  1. Add a self-assessment module for harness-induced variance.

The improved system will track its own run-to-run stability under different harness configurations. If it detects that avg@3 varies by more than 0.03 across harnesses, it will flag this and either switch to a more stable harness or adjust its internal parameters (e.g., reduce exploration noise) to compensate. This addresses the finding that harness choice mainly affects reliability, not peak performance.


What the improved AI system can do:

  • Diagnose its own failures in real time (direction vs. execution vs. recovery) and adapt its strategy accordingly.

  • Reuse experience selectively and safely, avoiding the performance degradation caused by blind context retention.

  • Recover quickly from regressions by always protecting its best verified state.

  • Transfer abstract lessons across tasks without carrying over irrelevant or harmful raw context.

  • Avoid brittle, evaluation-specific shortcuts and produce more generalizable solutions.

  • Maintain consistent performance across repeated runs, reducing the gap between average and best outcomes.

  • Tailor its approach to the specific bottleneck of each task family (e.g., more exploration for CUDA, more state-preservation for Model Development).

  • Automatically adjust its execution cadence (build frequency, verification steps) based on observed reliability.

Sources

Related papers