ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation

arXiv:2608.24103 · cs.AI · Submitted 2026-08-25 · Read on arXiv

cs.AI

Submitted: 2026-08-25

Updated: 2026-08-25

Comments: 26 pages, 12 figures, EMNLP 2026 Industry Track (Main paper)

Code: https://github.com/BloomBerry/agentic-canvas-editor

License: http://creativecommons.org/licenses/by/4.0/

The gist: Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two practical problems block reliable deployment: legacy document formats expose only flat,

Terminology

Abstract

Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two practical problems block reliable deployment: legacy document formats expose only flat, absolutely positioned elements, so agents must recompute coordinates and routinely break layouts; and design has no unique ground truth, so diff-against-reference metrics penalize valid-but-different outputs. We present ACE, an agentic canvas editor over a hierarchical scene-graph with a presentation-specialized action space (98 tools), paired with CARE, a content-aware router that feeds the agent only the relevant slice of each deck (avg. about 89% input-token reduction), and a self-correction loop driven by a ground-truth-free instruction-following (IF) judge whose natural-language critique is fed back as the next-turn instruction. With a fixed backbone, a scene-graph editor in a single turn already matches a same-backbone agentic HTML pipeline that iterates internally; adding self-correction lifts ACE significantly above it on instruction following (IF 4.23 vs. 3.81 on the full 94-task benchmark, paired p =.010, replicated by an out-of-loop judge) at 1.75 times the speed and about 44% lower cost. VQ means are statistically indistinguishable, but 26 blind raters prefer ACE overall (58.7% decisive win-rate) and prefer the self-corrected output 81% of the time; the ranking is invariant across three judge families, and out-of-loop judges retain two-thirds of the self-correction gain, bounding circularity. 66% of cases halt after one pass, and a strict-peak rollback removes every observed regression.

Sources

Related papers