TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images

summary

Video file (mp4)

The gist

Recent advances in text-to-image models have improved global realism, but persistent failures in local typography—such as malformed glyphs and broken strokes—remain significant perceptual quality

In short

This work introduces Text-in-Image Quality Assessment (TIQA), a new no-reference task to measure how well rendered text looks to humans, separate from whether the text is linguistically correct. The model ANTIQA was developed to predict this perceptual quality score using image features and edge maps. Results show ANTIQA accurately predicts human scores, proving that visual text fidelity is a crucial factor in overall image preference.

Key concepts

TIQA
A task designed to estimate the human-aligned perceptual quality of rendered text regions within an image. It focuses on typographic artifacts like malformed glyphs or broken strokes rather than checking if the words are spelled correctly.
ANTIQA
A lightweight model proposed to predict the TIQA score. It takes an input image and a Sobel edge map, processes them through convolutional blocks, and uses an Adaptive Pooling Block to output a single scalar quality score for the text region.
No-Reference IQA
Image Quality Assessment (IQA) without any reference image or ground truth quality score. TIQA is a specialization of this field, focusing specifically on identifying visual degradations related to typography in AI-generated images.
Semantic Correctness vs. Visual Fidelity
The distinction between whether text is linguistically correct and whether it looks visually good. TIQA aims to measure visual fidelity—how well the rendered characters look—independent of the text's actual meaning or grammatical accuracy.

Terminology used across episodes

This episode discusses

The paper

TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images · Read on arXiv

Kirill Koltsov, Aleksandr Gushchin, Anastasia Antsiferova, Dmitriy Vatolin

Lomonosov Moscow State University · ISP RAS Research Center for Trusted Artificial Intelligence

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images".

Jane: Recent advances in text-to-image models have improved global realism,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about the title and who came up with this paper, "TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images." It’s pretty descriptive, focusing on aligning the quality score with human perception of rendering artifacts.

Jane: Exactly. The authors are Kirill Koltsov, Aleksandr Gushchin, Anastasia Antsiferova, and Dmitriy Vatolin from Lomonosov Moscow State University and the ISP RAS Research Center for Trusted Artificial Intelligence in Russia.

Lu: Their focus on disentangling visual text quality from semantic correctness is really clever; it sets up a framework where we can evaluate the rendering process independently of linguistic accuracy.

Meng: The authors are introducing this no-reference task specifically to address those persistent failures in local typography that current evaluations simply don't capture well.

Lalam: It’s important because if we can build tools that predict these perceptual failures, we can start creating quality signals directly related to how text looks in an image, which is something we need for better training.

The paper's summary: Tom: So, what’s the core idea they’re presenting? Essentially, they formulate TIQA as a no-reference task that predicts a scalar quality score for a detected text region that matches human judgments of rendered-text fidelity.

Jane: That means we are asking an AI to predict how well the text looks visually—things like if the strokes are continuous or if the letters are correctly formed—without needing to know what the text is actually saying.

Lu: They’ve introduced two datasets, TIQACrops with 120k crops and TIQA-Images with one thousand five hundred full-frame images paired with human MOS scores for overall and text-only quality.

Meng: Having those specific datasets is crucial because it gives the model something concrete to train on that directly reflects what people actually rate in terms of text rendering artifacts.

Lalam: The paper emphasizes excluding semantic correctness from the goal, stating that semantic correctness is better measured by OCR models or VLM-based recognition methods.

The paper's improvements: Tom: The authors propose a specific model called ANTIQA, which they describe as a lightweight predictor with text-specific inductive biases designed to capture both fine-grained glyph details and global word structure.

Jane: Their architecture uses an input that combines the image crop with a Sobel edge map, feeding it into three resolution stages of repeated ConvB blocks separated by DownScale modules.

Lu: The use of Adaptive Pooling Block to produce a fixed-size scale embedding seems like it helps the network capture those specific visual features at different levels of abstraction effectively.

Meng: Their training objective is a mix of L = LMSE plus lambda Lrank, which aims to encourage both calibrated scores and correct relative preferences between different text regions.

Lalam: This mixed loss function seems really smart because it ensures the model not only gets the absolute score right but also understands how one piece of text quality compares to another in the same image.

Conclusion: Tom: So, wrapping up this discussion on "TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images," the main takeaway is that we have a dedicated task for assessing perceptual text artifacts independent of semantic correctness.

Jane: It really shows that capturing these visual failures—like broken strokes or irregular spacing—is a significant bottleneck in current text-to-image generation, even when the overall image looks realistic.

Lu: The results on TIQACrops show ANTIQA reaching a PLCC/SROCC of zero point nine four two/zero point nine three five, which demonstrates strong alignment with human judgments on crop level MOS, outperforming OCR confidence scores and VLM judges there.

Meng: On the larger TIQA-Images dataset, they found strong correlations with both overall and text-only MOS values, suggesting that improvements in text rendering quality often coincide with improvements in the overall perceived quality of the image.

Lalam: This work gives us a powerful control signal: we can use this score to filter candidates during generation or to guide closed-loop controls where low scores trigger re-rendering without changing the original prompt semantics.

Tom: It’s exciting because it validates that treating rendered text fidelity as a distinct evaluation target is necessary for improving the quality of text-heavy generations.

Jane: Yes, and we have to remember their limitation: the paper notes that visual plausibility systematically overstates text fidelity, meaning an image can look convincing while the embedded text itself remains degraded.

Lu: That points toward a future direction where we need new feature extractors specifically designed to capture typographic artifacts like kerning instability or stroke continuity that general metrics miss.

Meng: If we can build those specialized feature extractors, it opens the door for fine-tuning models directly on this loss function to improve rendering fidelity while keeping the prompt semantics fixed.

Lalam: It’s a strong foundation for improving our systems by providing a specific reward signal during sampling or training to optimize for high typographic plausibility.

More episodes

← Home