UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
cs.CV, cs.AI
Submitted: 2026-10-07
Updated: 2026-10-07
Code: https://github.com/LINs-lab/UltraText_Bench
Terminology
Sources
- TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
- LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
- PaddleOCR 3.0 Technical Report
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
- FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models
- VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
- TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images
- JoyType: A Robust Design for Multilingual Visual Text Creation
- Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
- BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation
- TextSculptor: Training and Benchmarking Scene Text Editing
- Efficient Generative Model Training via Embedded Representation Warmup
- Self-Adversarial One Step Generation via Condition Shifting
- OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
- GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation
- Aligning Text-to-Image Diffusion Models with Reward Backpropagation
- TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
- Visual Text Processing: A Comprehensive Review and Unified Evaluation
- Three-Body Scattering for Generative Modeling
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models