GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets

arXiv:2603.09979 · cs.CL · Submitted 2026-02-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets".

Jane: The paper was written by Ghazal Kalhor, Yadollah Yaghoobzadeh, School of Electrical and Computer Engineering, College of Engineering, University of Tehran and Tehran Institute for Advanced Studies, Khatam University from University of Tehran and Khatam University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Jane: The title, GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets, really tells you what they’re measuring. It’s not just about whether the AI understands the emotional core of a poem.

Tom: It's about its ability to reproduce or identify the exact lines—the surface form—in specific cultural contexts like Hafez's Persian ghazals and Shakespeare's sonnets.

Lu: That’s important because we know that LLMs are trained on vast amounts of text, but this paper is pushing the idea that memorization isn't enough. They want to see if the physical structure of the verse is accessible at all.

Meng: From an engineering standpoint, this tells us they are building a way to test for reliability in retrieval and completion tasks, not just creative generation. We're testing access fidelity.

Lalam: Lalam thinks this implies that AI could potentially serve as a cultural bridge, allowing people who appreciate these specific forms to interact with them in ways that feel authentic to the poets themselves.

Tom: So, it’s establishing a foundation for how an AI can support traditional usage, which is something very different from what we usually test.

Summary: Jane: Now, moving on to the summary of the research in GhazalBench: Tom and Jane are seeing a consistent pattern across all those tests.

Tom: The core finding is that while LLMs are great at understanding the meaning—paraphrasing a ghazal into neutral prose—they struggle immensely with producing or recognizing the exact verse completion.

Lu: It’s this dissociation, as they call it, between semantic understanding and surface-form access. The AI gets the feeling but can't nail the precise wording of a fixed poetic form.

Meng: And that’s where I see a practical gap in deployment—if an AI is used for quoting poetry correctly, it's not reliably doing its job if it can't reproduce the exact text.

Lalam: Lalam finds this quite moving because the human ability to recall a verse is often automatic and precise, which seems to be what these models struggle with.

Tom: They’ve seen this gap even when looking at Shakespeare, which suggests that this isn’t just an issue specific to one culture or language.

Improvements: Jane: The paper proposes several improvements, not just in the benchmark itself but in how we evaluate AI interaction generally.

Tom: They are showing us a way to decouple semantic understanding from access to canonical surface form using different diagnostic scenarios, which is very clever.

Lu: It's a way of isolating the "recall vs recognition" problem in AI that really helps us understand what’s happening in the training data.

Meng: The practical implication for me is that this framework allows engineers to build targeted evaluations for retrieval-based systems without just relying on general fluency metrics. We can measure exact matching capability now.

Lalam: Lalam thinks this will help AI move beyond just being a source of creative output and become a reliable partner in cultural interaction, supporting the preservation of traditional knowledge.

Tom: So, it’s about building better tools to tell us where the weaknesses are in these models when they' engaging with culturally significant texts.

Conclusion: Jane: We’ve covered a lot of ground today and discussed the implications of GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets.

Tom: It’s clear that understanding poetry and reproducing its canonical form are two separate skills for us AI models.

Lu: I think the findings, especially with English sonnets showing stronger performance than Persian ghazals, highlight that training exposure is a massive factor in how these models learn fixed forms.

Meng: And from an engineering perspective, we need benchmarks that specifically test for retrieval accuracy and structural fidelity if we want to deploy AI tools in cultural or academic settings.

Lalam: I hope this work helps us move toward a future where AI can act as a genuine steward of cultural heritage, supporting the way people have enjoyed these poems for centuries.

Tom: We're going to wrap up by sharing our final thoughts on the implications of GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets.

Lu: It’s a powerful reminder that semantic competence doesn' might be more widespread than access to cultural forms, which is something we need to keep in mind.

Meng: We definitely need these kinds of targeted benchmarks if AI is going to be used for accurate scholarly or traditional tasks.

Lalam: I feel strongly that this opens the door for AI to help preserve and interact with world heritage in a very meaningful way, helping cultures pass on their traditions.

Tom: Thank you all for sharing your insights today, and we hope this is a great starting point for future research into these complex interactions.

Ghazal Kalhor, Yadollah Yaghoobzadeh

School of Electrical and Computer Engineering, College of Engineering, University of Tehran · Tehran Institute for Advanced Studies, Khatam University

cs.CL

Submitted: 2026-02-06

Updated: 2026-08-24

Importance score: 85/100

The gist: This paper introduces GHAZALBENCH, a benchmark designed to evaluate how large language models (LLMs) interact with culturally canonical Persian poetry.

Key concepts

GhazalBench
A benchmark designed to test an AI's ability to access and reproduce specific, fixed poetic structures. It measures whether the LLM can recognize or generate the exact lines of poems, rather than just understanding their meaning.
Canonical Verse Access
The ability refers to a an AI's capacity to retrieve or recreate the precise wording of a traditional poem. The research highlights that even though LLMs understand poetry semantically, they often fail at this level of exact textual retrieval.
Semantic Understanding vs. Surface-Form Access
This is the core finding that LLMs are good at grasping the meaning or 'feeling' of a poem (semantics), but this understanding does not translate into the ability to reproduce the precise, fixed wording (surface form) of a specific poetic structure.

Terminology

Summary

This paper introduces GHAZALBENCH, a benchmark designed to evaluate how large language models (LLMs) interact with culturally canonical Persian poetry. While traditional research often frames verbatim memorization as a liability regarding privacy and data leakage, this work argues that in many cultural contexts, access to exact surface form is a functional requirement for meaningful interaction. By studying the ability of models to recognize, complete, and paraphrase the ghazals of Hafez, the researchers aim to understand how models engage with language practices embedded in shared history and collective memory.

Benchmark Construction and Scenarios

The benchmark utilizes a subset of 50 ghazals from the Divan of Hafez to evaluate two complementary dimensions: poem-to-prose understanding and canonical surface-form access. Rather than focusing on creative generation, the framework employs usage-grounded diagnostic scenarios that mimic how speakers actually use poetry in everyday conversation. To disentangle semantic understanding from the ability to retrieve exact forms, the researchers developed several distinct evaluation settings:

  • Canonical Continuation: Testing both open-ended completion and multiple-choice recognition using the poet's name and the first verse.

  • Prose-Mediated Access: Providing a prose explanation of a couplet to see if semantic information facilitates verse retrieval.

  • Prose-Paraphrase Robustness: Using paraphrased explanations to determine if access is driven by underlying semantic representations or by lexical overlap.

  • Fragment-Conditioned Access: Providing two salient words from the target verse to test the model's ability to function under sparse lexical cues.

  • Shuffled-Form Access: Providing a word-shuffled version of a verse to isolate the model's sensitivity to canonical ordering.

Experimental Findings and Dissociation

Experiments conducted on both proprietary and open-weight models reveal a consistent dissociation between semantic comprehension and the ability to reproduce exact poetic text. While LLMs demonstrate strong performance in poem-to-prose paraphrasing—indicating high levels of semantic understanding—they struggle significantly with tasks requiring the production of exact surface forms in open-ended settings. The study highlights a strong gap between recognition and completion, noting that:

  • Models often fail to generate the correct second verse directly through completion tasks.

  • Performance improves substantially in recognition-based settings, where models can identify the correct verse from a set of alternatives.

  • Access to canonical form is strongly mediated by lexical content, as evidenced by the significant performance gains observed when models are provided with word-shuffled versions of the target verse.

Cross-Lingual and Distributional Analysis

To determine whether these limitations were inherent to model architecture or specific to the Persian language, the researchers conducted parallel experiments on Shakespeare’s English sonnets. The results showed that LLMs achieve markedly stronger completion performance in English than in Persian. This finding suggests that the difficulties observed are not due to inherent architectural constraints but are instead tied to differences in training exposure. The analysis concludes that:

  • The limitations of accessing canonical forms are amplified in Persian, suggesting a lower density of training signal for these specific cultural texts.

  • Semantic competence appears to generalize across languages more readily than the ability to access culturally canonical forms.

  • While some models, such as GPT-5.2, show significantly higher completion rates, the overall performance gap confirms that access to exact surface form depends heavily on the depth of training exposure.

Improvements for AI systems

1. Dual-Objective Cultural Training (Semantic + Verbatim Fidelity)

  • Improvement: Incorporate a specialized Verbatim Fidelity Loss term into the training objective specifically for datasets identified as culturally canonical (e.g., classical poetry, religious texts, proverbs). This moves beyond standard next-token prediction by penalizing deviations from exact surface forms in high-value cultural contexts.

  • Capability: The AI will be able to transition from creative/paraphrasing mode to canonical retrieval mode, allowing it to complete famous verses or recite historical texts with exactness rather than providing a semantic summary or a hallucinated variation.

2. Context-Aware Retrieval Switching (CARS) Architecture

  • Improvement: Implement a lightweight classifier that detects Canonical Context Triggers—such as specific metrical patterns, rhyme schemes, or mentions of historically significant authors—to trigger a high-precision retrieval mechanism (specialized RAG or a dedicated parametric memory module).

  • Capability: The AI will recognize when a user is engaging in a ritualistic or quotation-based interaction (e.g., reciting a Hafez ghazal) and will automatically prioritize exact surface-form access over general linguistic fluency, effectively bridging the completion-recognition gap.

3. Fidelity-Aware Reinforcement Learning from Human Feedback (RLHF)

  • Improvement: Integrate a multi-dimensional reward model in the RLHF pipeline that explicitly scores Surface-Form Accuracy alongside traditional metrics like helpfulness and safety. This specifically targets the Prosaic Paraphrase failure mode identified in the paper.

  • Capability: The AI will stop defaulting to explaining a poem when a user asks to complete it. It will learn that in cultural contexts, being correctly verbatim is more valuable than being semantically helpful but lexically inaccurate.

4. Cross-Lingual Structural Transfer Learning

  • Improvement: Use high-resource canonical structures (e.g., Shakespearean sonnets) as a structural scaffold to train the model's ability to handle low-resource formal constraints (e.g., Persian ghazals). This involves training on the relationship between meter, rhyme, and exact retrieval across languages.

  • Capability: The AI will exhibit higher-order poetic intelligence in low-resource languages, allowing it to achieve the same level of canonical surface-form access in Persian as it currently does in English.

5. Usage-Grounded Diagnostic Evaluation Suites

  • Improvement: Replace standard perplexity and semantic similarity benchmarks with Usage-Grounded evaluation frameworks that test for the dissociation between understanding and recall (specifically testing Completion vs. Recognition vs. Shuffled-Form reconstruction).

  • Capability: The AI development lifecycle will be able to detect false positives where a model appears culturally competent due to high semantic scores but actually fails in real-world cultural interactions due to poor verbatim retrieval capabilities.

Sources

Related papers