EdiTikZ: Scientific Figure Editing from Revision Trajectories
cs.AI, cs.CL, cs.CV
Submitted: 2026-09-01
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Vision-language models (VLMs) have shown strong performance in generating scientific figures from text or images.
Terminology
Abstract
Vision-language models (VLMs) have shown strong performance in generating scientific figures from text or images. However, producing publication-ready figures requires iterative refinement, making scientific figure editing an important yet largely unexplored task. Existing approaches rely on costly proprietary agentic systems, focus primarily on evaluation, or construct training supervision from synthetically generated edits. Instead, we leverage naturally occurring scientific revision and development trajectories as a scalable source of supervision. To this end, we introduce DaEdiTikZ, the first large-scale dataset of revision-derived scientific figure edits, constructed by mining 391K plausible TikZ edit pairs from arXiv, GitHub, and TeX SE and inferring 781K directed edit instructions with a VLM conditioned on rendered figures and TikZ code. We further introduce DaEdiTikZ-Bench, a human-refined benchmark with 790 instances, and train two compact Qwen3.5-based EdiTikZ models (4B and 9B) by jointly learning reconstruction and editing, followed by reinforcement learning (RL) with complementary rewards for rendered fidelity and edit application. Automatic evaluation places our 9B model above all tested baselines, while human evaluation with 9 annotators and 4,320 ratings places it above GPT-5.6-Sol and on par with Gemini-3.1-Pro. Under severe out-of-distribution shifts, it remains competitive with GPT-5.6-Sol near its 2K training sequence-length regime. Models and datasets will be released.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
- Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
- GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing
- Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning
- AutoFigure-Edit: Generating Editable Scientific Illustration
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
- Decoupled Weight Decay Regularization
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
- DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing
- Multimodal Chain-of-Thought Reasoning in Language Models
- Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
- PaperBanana: Automating Academic Illustration for AI Scientists
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection