Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
cs.CV, cs.AI
Submitted: 2026-09-28
Updated: 2026-09-28
Project page: https://ir4rl.github.io
Terminology
Sources
- Qwen3-VL Technical Report
- TikZero: Zero-Shot Text-Guided Graphics Program Synthesis
- VisRefiner: Learning from Visual Differences for Screenshot-to-Code Generation
- Process-Verified Reinforcement Learning for Theorem Proving via Lean
- Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback
- Visual-ERM: Reward Modeling for Visual Equivalence
- DINOv2: Learning Robust Visual Features without Supervision
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning
- Solving math word problems with process- and outcome-based feedback
- Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
- Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
- OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens
- UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
- Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models