ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation
cs.AI
Submitted: 2026-08-25
Updated: 2026-08-25
Comments: 26 pages, 12 figures, EMNLP 2026 Industry Track (Main paper)
Code: https://github.com/BloomBerry/agentic-canvas-editor
License: http://creativecommons.org/licenses/by/4.0/
The gist: Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two practical problems block reliable deployment: legacy document formats expose only flat,
Terminology
Abstract
Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two practical problems block reliable deployment: legacy document formats expose only flat, absolutely positioned elements, so agents must recompute coordinates and routinely break layouts; and design has no unique ground truth, so diff-against-reference metrics penalize valid-but-different outputs. We present ACE, an agentic canvas editor over a hierarchical scene-graph with a presentation-specialized action space (98 tools), paired with CARE, a content-aware router that feeds the agent only the relevant slice of each deck (avg. about 89% input-token reduction), and a self-correction loop driven by a ground-truth-free instruction-following (IF) judge whose natural-language critique is fed back as the next-turn instruction. With a fixed backbone, a scene-graph editor in a single turn already matches a same-backbone agentic HTML pipeline that iterates internally; adding self-correction lifts ACE significantly above it on instruction following (IF 4.23 vs. 3.81 on the full 94-task benchmark, paired p =.010, replicated by an out-of-loop judge) at 1.75 times the speed and about 44% lower cost. VQ means are statistically indistinguishable, but 26 blind raters prefer ACE overall (58.7% decisive win-rate) and prefer the self-corrected output 81% of the time; the ranking is invariant across three judge families, and out-of-loop judges retain two-thirds of the self-correction gain, bounding circularity. 66% of cases halt after one pass, and a strict-peak rollback removes every observed regression.
Sources
- AutoPresent: Designing Structured Visuals from Scratch
- CANVAS: A Benchmark for Vision-Language Models on Tool-Based User Interface Design
- Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation
- Revision Matters: Generative Design Guided by Revision Edits
- PPTArena: A Benchmark for PowerPoint Editing
- AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Reflexion: Language Agents with Verbal Reinforcement Learning
- SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design
- ReAct: Synergizing Reasoning and Acting in Language Models
- VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection