Enhanced Video Text Editing with Trajectory-Aligned Glyph Rendering
cs.CV
Submitted: 2026-09-28
Updated: 2026-10-06
Terminology
Sources
- Qwen2.5-VL Technical Report
- Kling-Omni Technical Report
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing
- JoyType: A Robust Design for Multilingual Visual Text Creation
- PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System
- Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance
- InnoText: A Unified Model for Visual Text Generation and Editing
- Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering
- GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation
- CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder
- OpenAI GPT-5 System Card
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- HunyuanVideo 1.5 Technical Report
- Seedance 2.0: Advancing Video Generation for World Complexity
- AnyText2: Visual Text Generation and Editing With Customizable Attributes
- Wan: Open and Advanced Large-Scale Video Generative Models
- SteerVTE: Seamless Video Text Editing with Style and Glyph Control
- PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models