AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

arXiv:2608.13560 · cs.CV, cs.AI, cs.CL · Submitted 2026-08-13 · Read on arXiv

Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, Xiaotong Li

Meituan · MBZUAI · Huazhong University of Science and Technology · Peking University · Tsinghua University · The Chinese University of Hong Kong · Shanghai Jiao Tong University

cs.CV, cs.AI, cs.CL

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Tech Report. Code at: https://github.com/Yaxin9Luo/AutoDesign

Code: https://github.com/Yaxin9Luo/AutoDesign

Project page: https://autodesign.designanything.ai

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper addresses the challenge of transforming multimodal sources into condensed and structured media outputs (e.g., posters, slides, webpages, videos), which the authors conceptualize as "a

Terminology

Summary

The paper addresses the challenge of transforming multimodal sources into condensed and structured media outputs (e.g., posters, slides, webpages, videos), which the authors conceptualize as a long-horizon agentic process centered on a model-harness system. The key gap identified is that "while an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability. The authors note that unlike human creators who continuously accumulate knowledge from successful revisions and failures, such systems treat individual human-aligned feedback as transient signals rather than reusable design knowledge."

The paper presents AutoDesign, a framework that aligns with human design priors, where a meta-harness optimizer guides a code agent to recursively improve harness based on rollout feedback. The framework operates through two nested loops:

  • Inner loop (design harness): transforms the source context into an editable output and iteratively revises it under critic feedback

  • Outer loop (meta-harness): "optimizes the design harness across tasks. It first grounds human preferences by initializing an evaluator from annotated reference artifacts. Given this human-aligned evaluator, the meta-harness aggregates rollouts and evaluation scores across tasks to identify recurrent failures, and directs a coding agent as the optimizer to propose a bounded update to the current design harness at each iteration."

The design harness is decomposed into five functional components: Context and Memory, Tools and Specifications, Execution Runtime, Orchestration, and Evaluation and Feedback. The optimization objective is formalized as H⋆ = arg max H J(H), where J(H) measures the expected quality of the artifacts produced by the design harness.

The outer loop proceeds through four stages per iteration:

  1. Rollout: the current design harness Ht is executed on a training task set Dtrain

  2. Evaluation: an evaluator coding agent with reference artifacts annotated by humans along seven quality dimensions implements the evaluator Rmeta

  3. Update proposal: the optimizer takes the current harness Ht, the collected trajectories τt, their evaluation scores st, and the optimization record L as input, and produces a candidate updated harness

  4. Acceptance gate: A candidate is accepted only when its performance on Dtrain improves and its performance on Ddev does not decline

The acceptance condition is: Accept(H′t+1) ⇐⇒ Jtrain(H′t+1) > Jtrain(Ht) ∧ Jdev(H′t+1) ≥ Jdev(Ht). The paper emphasizes that Each outer-loop iteration is restricted to exactly one of the five harness components to keep credit assignment interpretable.

The optimized system, DesignHarness, consists of four main stages:

  1. Paper Ingestion: transforms the input source x and the design context c into a structured, provenance-aware context including document metadata and section outline, identifies key passages supporting the main claims, and records figures and tables together with their source locations

  2. Artifact Generation and Revision: The designer module is implemented as a coding agent that generates or revises the artifact from the ingested source context with the artifact remains as editable HTML files throughout refinement, allowing revisions to be implemented as localized code edits without requiring regeneration of the entire output

  3. Validation: the candidate artifact yk produced by the designer is examined by the rule-based validator, which applies a set of deterministic blocking checks covering unsafe or missing assets, broken provenance links between incorporated materials and their sources, severe overflow or overlap, and violations of the required typographic and layout constraints

  4. Finalization: applies the remaining post-processing, such as final rendering adjustments, mathematical typesetting, and inlining of referenced assets, to produce a self-contained output

The system permits at most K = 12 refinement attempts and uses a critic VLM that assesses rendered properties of the candidate, including compliance with the design context, layout, readability, and aesthetics.

The paper introduces PosterBench, a comprehensive evaluation protocol for paper-to-poster evaluation containing a 100-paper Main Track and PosterBench-mini, a shared 10-paper subset. The papers span five disciplines: AI/ML, biomedicine and health, climate and earth environment, economics and policy, and physics and astronomy.

The evaluation uses a seven-dimensional rubric with weights: Faithfulness (10), Coverage (10), Density (15), Visual Evidence (10), Layout (20), Readability (25), and Aesthetics (10). The final score applies the strictest active record-level ceiling before benchmark averaging with protected gates including severe layout damage, insufficient presentation viability, confirmed visible failures, and protected render-integrity violations.

On the PosterBench Main Track, AutoDesign achieves the highest score of 78.32, surpassing the closed-source commercial system Claude Design by 7.45 points. Under the matched Claude Code and Claude 4.8 configuration, AutoDesign scores 78.32, exceeding Claude Design by 7.45 points and OpenDesign by 8.87 points.

Across seven controlled code agent–model configurations, integrating the learned DesignHarness consistently improves performance, increasing the average PosterBench Score from 54.99 to 67.39 (+12.4%). The gains range from +5.59 (Codex/GPT-5.5: 75.87 to 81.46) to +19.56 (DeepSeek V4 Pro with Claude Code: 34.73 to 54.29).

"The observed Pareto frontier runs from LongCat-2.0 (55.13 at 0.27 per poster), through Doubao Seed 2.1 Pro (71.83 at 2.75) and Claude 4.8 (74.56 at 7.63), to GPT-5.5 (81.46 at 10.02). Doubao reaches 88% of the GPT-5.5 score at 27% of its cost."

In a fully autonomous long-horizon loop, it executes 253 tool calls and 11 editing turns within 40 minutes for under 3, reaching average conference-poster quality in human evaluation. In the system-blind study, eleven reviewers submitted 936 responses (933 ranking judgments and three skips); AutoDesign's probability of beating a randomly sampled alternative was 64.0%. The head-to-head results show AutoDesign preferred in 44% vs. Claude Code, 51% vs. OpenDesign, and 55% vs. Claude Design of pairwise judgments.

The two measurements have a positive, albeit imperfect, association (r = 0.34; a paper-cluster bootstrap gives a 95% interval of [0.22, 0.44]). The probability that the benchmark-preferred poster agrees with human decisions rises from 51.9% for 0–3-point gaps to 74.4% for gaps of at least 20 points.

The paper reports that Autonomous optimization rises from 49.00 to an 80.88 plateau; redirected search reaches 88.39 for a representative paper. Through "7 days of evolving traces, it invokes 224 subagents, records at least 123 recursive iterations, accumulating 54 harness updates that recursively convert human-designed reference artifacts, rollout traces, rendering diagnostics, and evaluator feedback into persistent design priors."

The paper notes that AutoDesign is presently validated for academic paper-to-poster generation, but the underlying agentic design pattern is not tied to a single input or output medium. Pilot artifacts demonstrate paper-to-slide, paper-to-webpage, and paper-to-conference-video outputs. The authors state that Each medium needs source–output data, an evaluator, a rendering and validation gate, and an objective tailored to its communication setting.

The paper lists four main contributions:

  1. AutoDesign, a meta-harness optimization framework that turns a static design harness into a recursively improving system for human-aligned multimodal design

  2. DesignHarness, the executable academic paper-to-poster system evolved by AutoDesign

  3. PosterBench, a comprehensive evaluation protocol for paper-to-poster evaluation

  4. Demonstration that under a fully autonomous long-horizon agentic loop, DesignHarness produces human-level academic posters

Improvements for AI systems

Improvements to AI Systems:

  1. Recursive Self-Improvement via Meta-Harness Optimization
  • Implement an outer loop that continuously evaluates an agent's own workflow (not just outputs) using a human-aligned evaluator, then proposes bounded updates to the agent's internal tools, orchestration, or feedback mechanisms.

  • The improved system can autonomously refine its own design/creation pipeline over multiple tasks, accumulating reusable design priors from failures and successes, rather than treating each task as isolated.

  1. Human-Preference Grounding from Reference Artifacts
  • Initialize an evaluator module by learning from annotated human-designed examples (e.g., posters, slides) across multiple quality dimensions (e.g., faithfulness, layout, readability, aesthetics).

  • The improved system can generate outputs that better match human aesthetic and structural preferences from the start, reducing the need for manual prompt engineering or post-hoc corrections.

  1. Structured, Provenance-Aware Context Ingestion
  • Decompose input sources (e.g., papers) into structured metadata, key claims, supporting passages, and visual assets with source links, before generation.

  • The improved system can produce outputs with verifiable traceability (every claim/visual linked to its source), reducing hallucination and enabling automatic fact-checking.

  1. Localized, Editable Artifact Representation
  • Keep outputs as editable code (e.g., HTML) throughout refinement, allowing targeted revisions (e.g., fixing a layout issue) without regenerating the entire artifact.

  • The improved system can perform surgical edits, saving compute and enabling faster iteration, especially for long-form or complex outputs.

  1. Deterministic Validation Gates Before Aesthetic Critique
  • Apply rule-based blocking checks (e.g., missing assets, broken links, overflow, layout violations) before invoking a visual-language-model critic.

  • The improved system avoids wasting expensive model calls on fundamentally broken outputs, ensuring that only structurally sound candidates receive aesthetic feedback.

  1. Bounded, Interpretable Update Proposals
  • Restrict each optimization iteration to modifying exactly one component of the agent's harness (e.g., only the evaluator or only the orchestration), with an acceptance gate requiring improvement on training tasks and no regression on held-out tasks.

  • The improved system can be safely and incrementally improved without catastrophic forgetting or unpredictable behavior changes.

  1. Cost-Aware Pareto Optimization
  • Track performance vs. cost across different model–agent configurations, and allow the meta-harness to select or recommend configurations based on user budget constraints.

  • The improved system can automatically trade off quality and cost (e.g., achieving 88% of top-tier quality at 27% of the cost), making high-quality generation accessible on limited budgets.

  1. Cross-Medium Transferable Design Patterns
  • Use the same meta-harness framework (ingestion → generation → validation → finalization) for different output types (posters, slides, webpages, videos) by swapping only the medium-specific evaluator, rendering gate, and objective.

  • The improved system can adapt its learned design priors across media, reducing the need to retrain from scratch for each new output format.

  1. Autonomous Long-Horizon Operation with Minimal Human Oversight
  • Enable the system to execute hundreds of tool calls and multiple editing turns over extended periods (e.g., 40 minutes) without human intervention, while maintaining quality.

  • The improved system can handle complex, multi-stage creative tasks (e.g., conference poster design) end-to-end, freeing human experts for higher-level decisions.

  1. Benchmark-Human Alignment Calibration
  • Use a benchmark with a strict ceiling and protected gates (e.g., severe layout damage) to filter out catastrophic failures, and calibrate benchmark scores against human preferences to ensure they correlate meaningfully.

  • The improved system can be evaluated more reliably, with benchmark scores that predict human satisfaction, enabling trustworthy automated quality assessment.


What the Improved AI System Can Do:

  • Automatically design conference posters, slides, webpages, or videos from academic papers, matching human-level quality and aesthetics.

  • Continuously improve its own design workflow across tasks, learning from past failures and human preferences without manual reprogramming.

  • Produce outputs with fully traceable provenance, editable code, and structural integrity, while optimizing for both quality and computational cost.

  • Operate autonomously for long-horizon tasks, making hundreds of tool calls and iterative refinements, with minimal human oversight.

Abstract

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability. In this paper, we present AutoDesign, a framework that aligns with human design priors, where a meta-harness optimizer guides a code agent to recursively improve harness based on rollout feedback. To instantiate and evaluate this framework, we focus on the academic paper-to-poster generation task and introduce PosterBench, comprising a 100-paper Main Track spanning five disciplines and PosterBench-mini, a shared 10-paper subset for controlled evaluation. On the PosterBench Main Track, AutoDesign achieves the highest score of 78.32, surpassing the closed-source commercial system Claude Design by 7.45 points. Across seven controlled code-agent-model configurations, integrating the learned DesignHarness consistently improves performance, increasing the average PosterBench Score from 54.99 to 67.39 (+12.4%). In a fully autonomous long-horizon loop, it executes 253 tool calls and 11 editing turns within 40 minutes for under 3, reaching average conference-poster quality in human evaluation. A system-blind human study further demonstrates that AutoDesign achieves the highest human preference among evaluated systems.

Sources

Related papers