Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence
cs.CL
Submitted: 2026-06-14
Updated: 2026-08-30
Comments: Accepted by TMLR 2026
Code: https://github.com/lukas-blecher/LaTeX-OCR
License: http://creativecommons.org/licenses/by/4.0/
The gist: While Large Language Models (LLMs) have substantially advanced text-to-code generation, many real programming tasks specify intent through visual artifacts such as screenshots, charts, and videos.
Terminology
Abstract
While Large Language Models (LLMs) have substantially advanced text-to-code generation, many real programming tasks specify intent through visual artifacts such as screenshots, charts, and videos. These tasks require models to connect visual perception to executable programs, as correctness depends not only on syntax but also on layout, data semantics, and domain-specific constraints that apply after execution. This survey reviews Multimodal Code Intelligence, covering systems that generate, edit, refine, or reason with code under visually grounded inputs and outputs. We first formulate the field by the role that code plays in each task, distinguishing code as a rendered artifact, an editable structure, an intermediate reasoning trace, or an executable tool interface. Then we organize benchmarks and methods into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Frameworks. This taxonomy connects artifact-generation problems to agentic and unified settings and allows us to compare how different tasks treat evidence of correctness. Across the literature, we argue that reliable evaluation requires evidence about semantics and interaction beyond visual fidelity. Looking ahead, future research may benefit from four verification-centered directions. Multi-signal validation can combine complementary evidence of correctness, multi-state verification can test behavior across execution trajectories, cross-task transfer testing can probe reusable visual-code skills, and verifiable agent traces can reveal whether agent actions are grounded in visual evidence. Together, these directions may move this field from single-output imitation toward evidence-grounded executable systems. An ongoing project and resources are available on GitHub.
Sources
- Program Synthesis with Large Language Models
- Query2CAD: Generating CAD models using natural language queries
- Qwen2.5-VL Technical Report
- StarFlow: Generating Structured Workflow Outputs From Sketch Images
- AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
- Multilingual Multimodal Software Developer for Code Generation
- Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
- CADReview: Automatically Reviewing CAD Programs with Error Detection and Correction
- Generative Interfaces for Language Models
- RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
- Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
- ChartEditor: A Reinforcement Learning Framework for Robust Chart Editing
- Evaluating Large Language Models Trained on Code
- InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
- AI4Research: A Survey of Artificial Intelligence for Scientific Research
- Ocean-OCR: Towards General OCR Application via a Vision-Language Model
- Logics-Parsing Technical Report
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- Code2Video: A Code-centric Paradigm for Educational Video Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering