TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images

arXiv:2603.07119 · cs.CV · Submitted 2026-03-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images".

Jane: Recent advances in text-to-image models have improved global realism,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about the title and who came up with this paper, "TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images." It’s pretty descriptive, focusing on aligning the quality score with human perception of rendering artifacts.

Jane: Exactly. The authors are Kirill Koltsov, Aleksandr Gushchin, Anastasia Antsiferova, and Dmitriy Vatolin from Lomonosov Moscow State University and the ISP RAS Research Center for Trusted Artificial Intelligence in Russia.

Lu: Their focus on disentangling visual text quality from semantic correctness is really clever; it sets up a framework where we can evaluate the rendering process independently of linguistic accuracy.

Meng: The authors are introducing this no-reference task specifically to address those persistent failures in local typography that current evaluations simply don't capture well.

Lalam: It’s important because if we can build tools that predict these perceptual failures, we can start creating quality signals directly related to how text looks in an image, which is something we need for better training.

The paper's summary: Tom: So, what’s the core idea they’re presenting? Essentially, they formulate TIQA as a no-reference task that predicts a scalar quality score for a detected text region that matches human judgments of rendered-text fidelity.

Jane: That means we are asking an AI to predict how well the text looks visually—things like if the strokes are continuous or if the letters are correctly formed—without needing to know what the text is actually saying.

Lu: They’ve introduced two datasets, TIQACrops with 120k crops and TIQA-Images with one thousand five hundred full-frame images paired with human MOS scores for overall and text-only quality.

Meng: Having those specific datasets is crucial because it gives the model something concrete to train on that directly reflects what people actually rate in terms of text rendering artifacts.

Lalam: The paper emphasizes excluding semantic correctness from the goal, stating that semantic correctness is better measured by OCR models or VLM-based recognition methods.

The paper's improvements: Tom: The authors propose a specific model called ANTIQA, which they describe as a lightweight predictor with text-specific inductive biases designed to capture both fine-grained glyph details and global word structure.

Jane: Their architecture uses an input that combines the image crop with a Sobel edge map, feeding it into three resolution stages of repeated ConvB blocks separated by DownScale modules.

Lu: The use of Adaptive Pooling Block to produce a fixed-size scale embedding seems like it helps the network capture those specific visual features at different levels of abstraction effectively.

Meng: Their training objective is a mix of L = LMSE plus lambda Lrank, which aims to encourage both calibrated scores and correct relative preferences between different text regions.

Lalam: This mixed loss function seems really smart because it ensures the model not only gets the absolute score right but also understands how one piece of text quality compares to another in the same image.

Conclusion: Tom: So, wrapping up this discussion on "TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images," the main takeaway is that we have a dedicated task for assessing perceptual text artifacts independent of semantic correctness.

Jane: It really shows that capturing these visual failures—like broken strokes or irregular spacing—is a significant bottleneck in current text-to-image generation, even when the overall image looks realistic.

Lu: The results on TIQACrops show ANTIQA reaching a PLCC/SROCC of zero point nine four two/zero point nine three five, which demonstrates strong alignment with human judgments on crop level MOS, outperforming OCR confidence scores and VLM judges there.

Meng: On the larger TIQA-Images dataset, they found strong correlations with both overall and text-only MOS values, suggesting that improvements in text rendering quality often coincide with improvements in the overall perceived quality of the image.

Lalam: This work gives us a powerful control signal: we can use this score to filter candidates during generation or to guide closed-loop controls where low scores trigger re-rendering without changing the original prompt semantics.

Tom: It’s exciting because it validates that treating rendered text fidelity as a distinct evaluation target is necessary for improving the quality of text-heavy generations.

Jane: Yes, and we have to remember their limitation: the paper notes that visual plausibility systematically overstates text fidelity, meaning an image can look convincing while the embedded text itself remains degraded.

Lu: That points toward a future direction where we need new feature extractors specifically designed to capture typographic artifacts like kerning instability or stroke continuity that general metrics miss.

Meng: If we can build those specialized feature extractors, it opens the door for fine-tuning models directly on this loss function to improve rendering fidelity while keeping the prompt semantics fixed.

Lalam: It’s a strong foundation for improving our systems by providing a specific reward signal during sampling or training to optimize for high typographic plausibility.

Kirill Koltsov, Aleksandr Gushchin, Anastasia Antsiferova, Dmitriy Vatolin

Lomonosov Moscow State University · ISP RAS Research Center for Trusted Artificial Intelligence

cs.CV

Submitted: 2026-03-07

Updated: 2026-09-29

Code: https://github.com/JaidedAI/EasyOCR

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Recent advances in text-to-image models have improved global realism, but persistent failures in local typography—such as malformed glyphs and broken strokes—remain significant perceptual quality

Key concepts

TIQA
A task designed to estimate the human-aligned perceptual quality of rendered text regions within an image. It focuses on typographic artifacts like malformed glyphs or broken strokes rather than checking if the words are spelled correctly.
ANTIQA
A lightweight model proposed to predict the TIQA score. It takes an input image and a Sobel edge map, processes them through convolutional blocks, and uses an Adaptive Pooling Block to output a single scalar quality score for the text region.
No-Reference IQA
Image Quality Assessment (IQA) without any reference image or ground truth quality score. TIQA is a specialization of this field, focusing specifically on identifying visual degradations related to typography in AI-generated images.
Semantic Correctness vs. Visual Fidelity
The distinction between whether text is linguistically correct and whether it looks visually good. TIQA aims to measure visual fidelity—how well the rendered characters look—independent of the text's actual meaning or grammatical accuracy.

Terminology

Summary

Recent advances in text-to-image models have improved global realism, but persistent failures in local typography—such as malformed glyphs and broken strokes—remain significant perceptual quality issues that current evaluations fail to capture. This work introduces Text-in-Image Quality Assessment (TIQA), a no-reference task designed to estimate a human-aligned perceptual quality score for rendered text regions while disentangling visual text quality from semantic correctness.

The gist

TIQA is the task of assessing perceptual failures rather than semantic correctness, predicting a scalar score for a detected text region that matches human judgments of rendered-text fidelity, independent of whether the string is linguistically correct.

New Task Formulation and Datasets

The paper introduces TIQA as a specialization of no-reference image quality assessment (IQA) specifically targeting typographic artifacts. The core contribution involves two datasets:

  1. TIQA-Crops: This dataset contains 120k text crops from 36k AI-generated images, annotated with mean opinion scores (MOS) and proxy labels for pretraining.

  2. TIQA-Images: This dataset consists of 1,500 full-frame images from recent generators, each paired with overall quality and text-only quality subjective scores (OQ-MOS and TQ-MOS).

The Proposed Model (ANTIQA)

The authors developed ANTIQA, a lightweight predictor with text-specific inductive biases. The architecture is designed to capture fine-grained glyph details and global word-level structure. Key architectural components include:

(i) Input:

ANTIQA maps an input batch I ∈ R B×2×H ×W (grayscale concatenated with a Sobel edge map) to a scalar quality score per sample.

(ii) Backbone:

The network utilizes a three resolution stages of repeated ConvB blocks separated by two DownScale modules that halve spatial size and double the channel count (64 → 128 → 256).

(iii) Feature Extraction:

Adaptive Pooling Block (APB) produces a fixed-size scale embedding.

(iv) Training Objective:

The model is trained using a mixed objective: L = LMSE + λ Lrank, which encourages both calibrated scores and correct relative preferences.

Analysis and Benchmarking

The performance of ANTIQA is evaluated against four baseline families, including OCR confidence scores (e.g., PaddleOCR), VLM-based judges (e.g., Qwen3-VL), generic no-reference IQA metrics (TOPIQ), and general supervised models (ResNet50, ViT).

(i) Crop-Level Results:

On TIQA-Crops, ANTIQA achieves strong alignment with human judgments, reaching PLCC/SROCC of 0.942/0.935. It outperforms OCR confidence and VLM judges on crop-level MOS, demonstrating that recognizing text is not sufficient: the model must also capture visual degradations specific to rendered glyphs (e.g., stroke breaks, bleeding).

(ii) Image-Level Results:

On TIQA-Images, ANTIQA achieves strong correlations with both overall and text-only MOS values. Notably, the OQ-TQ coupling is not merely driven by differences between generators but persists within fixed generator–prompt settings, suggesting that seed-level improvements in text rendering quality tend to coincide with seed-level improvements in overall perceived quality.

Downstream Applications

TIQA models provide a control signal for various generation pipelines:

  1. Reranking and filtering: Using predicted scores to rank multiple candidates per prompt by predicted text quality, or apply accept/reject thresholds.

  2. Quality-aware routing: TIQA can gate OCR outputs (accept vs. abstain), trigger re-generation/re-rendering, when synthetic artifacts cause failures in OCR or VLM systems.

  3. AI-image detection: TIQA scores can be fused with general real-vs-AI detectors to provide an additional text-specific signal.

  4. Guidance for T2I models: TIQA can be used as a reward for selection among samples, or as an auxiliary objective during sampling/training to improve rendered text while keeping prompt semantics fixed.

The findings establish perceptual text quality as a distinct evaluation target, showing that in text-heavy generation, rendered text is a major driver of overall human preference. The results demonstrate that ANTIQA consistently selects higher-MOS images and serves as a strong drop-in signal for filtering and reranking. Furthermore, the analysis shows that visual plausibility systematically overstates text fidelity: images can look convincing while the embedded text remains degraded. This highlights typography as an important unresolved bottleneck in modern text-to-image generation.

Improvements for AI systems

Here are specific improvements for AI systems based on the TIQA (Text-in-Image Quality Assessment) framework and ANTIQA model:

  1. The proposed AntiQA model can be integrated into a generation-time filtering pipeline to select the highest quality candidate images from a set of generations (e.g., in a best-of-K selection).

  2. AI systems can use the TIQA score to guide closed-loop control, where low scores trigger regeneration or reranking of the text/layout without changing the core prompt semantics.

  3. OCR and VLM reasoning pipelines can be made more robust by using TIQA as a gate: if a text region receives a low TIQA score (indicating likely malformed glyphs), the system can abstain from OCR/VLM interpretation for that specific region, falling back to simpler, more reliable methods or triggering an internal re-rendering loop.

  4. AI models can be fine-tuned using the TIQA loss function (combining MSE and a ranking loss) during training to directly improve the perceptual rendering fidelity of generated text while keeping prompt semantics fixed.

  5. A new text-specific feature extractor can be trained to capture typographic artifacts (glyph topology, stroke continuity, kerning instability) that are currently missed by general image quality metrics (like FID or BRISQUE), leading to more human-aligned quality scores for AI-generated text.

  6. Image generation systems can utilize TIQA feedback as a reward signal during sampling or training to optimize the output specifically for high typographic plausibility, moving beyond mere semantic correctness.

Abstract

Recent text-to-image models have improved global realism, but text rendering remains a persistent failure mode: images may look convincing overall, yet local typography often contains malformed glyphs, broken strokes, irregular spacing, and other artifacts that humans heavily penalize. We formulate Text-in-Image Quality Assessment (TIQA), a no-reference task that estimates a human-aligned perceptual quality score for detected text regions while disentangling visual text quality from semantic correctness. To support this setting, we introduce two datasets. TIQA-Crops contains 120k text crops from 36k AI-generated images produced by 12 generators, with 10k mean-opinion-score (MOS) labels and 110k proxy labels for pretraining. TIQA-Images contains 1,500 text-heavy images from 10 recent generators, including proprietary systems, with paired overall-quality and text-quality subjective scores. We also propose ANTIQA, a lightweight predictor with text-specific inductive biases. Across crop-level and image-level evaluations, ANTIQA achieves the best alignment with human judgments, reaching PLCC/SROCC of 0.942/0.935 on TIQA-Crops and 0.842/0.837 for text-quality MOS on unseen generators in TIQA-Images. In best-of-5 AI-generated image ranking, ANTIQA improves the text quality of the selected image by 0.36 MOS (14%), demonstrating utility for benchmarking, filtering, and generation-time selection. Together, these findings establish perceptual text quality as a distinct evaluation target for modern text-to-image generation.

Sources

Related papers