Figures as Programs: Recursive Generation of Editable Scientific Figures
cs.AI, cs.GR
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/yepengliu/FigTree
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Scientific methodology figures are essential for communicating complex methods clearly, yet creating them remains labor-intensive and typically requires multiple rounds of refinement.
Terminology
Abstract
Scientific methodology figures are essential for communicating complex methods clearly, yet creating them remains labor-intensive and typically requires multiple rounds of refinement. Recent image-generation models can synthesize visually appealing raster figures, but producing a human-satisfactory result in a single generation step remains difficult. Moreover, precise edits to raster figures are challenging for both humans and models. We formulate scientific figure generation as recursive SVG program construction and propose FigTree, a multi-agent system that automatically transforms a scientific paper into a structured vector figure. FigTree grounds figure content in the source paper, decomposes a figure into a hierarchy of local regions, generates each region as a short SVG program, and assembles the resulting fragments. A render-critic refinement loop jointly inspects the rendered figure and its underlying program, enabling visual defects to be traced to specific statements and accurately repaired. We conduct extensive evaluations of FigTree on figure quality and editability, showing that FigTree produces high-quality figures, while also enabling more effective editing than existing raster-based methods.
Sources
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- AI4Research: A Survey of Artificial Intelligence for Scientific Research
- MARS: Modular Agent with Reflective Search for Automated AI Research
- Towards Execution-Grounded Automated AI Research
- Qwen2.5-VL Technical Report
- GPT-4o System Card
- Gemini: A Family of Highly Capable Multimodal Models
- PaperBanana: Automating Academic Illustration for AI Scientists
- ConvexBench: Can LLMs Recognize Convex Functions?
- Qwen-Image-2.0 Technical Report
- Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
- AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations
- Automatic Method Illustration Generation for AI Scientific Papers via Drawing Middleware Creation, Evolution, and Orchestration
- Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations
- AutoFigure-Edit: Generating Editable Scientific Illustration
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection