GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets
summary
The gist
This paper introduces GHAZALBENCH, a benchmark designed to evaluate how large language models (LLMs) interact with culturally canonical Persian poetry.
In short
The discussion of 'GhazalBench' explores a study testing Large Language Models (LLMs) on Persian Ghazals and Shakespearean Sonnets. The core finding is that while LLMs can understand the meaning of poetry, they struggle to reproduce or recognize the exact canonical verse form. The hosts conclude that semantic understanding is separate from surface-form access, suggesting new benchmarks are needed for AI to serve as a reliable cultural steward.
Key concepts
- GhazalBench
- A benchmark designed to test an AI's ability to access and reproduce specific, fixed poetic structures. It measures whether the LLM can recognize or generate the exact lines of poems, rather than just understanding their meaning.
- Canonical Verse Access
- The ability refers to a an AI's capacity to retrieve or recreate the precise wording of a traditional poem. The research highlights that even though LLMs understand poetry semantically, they often fail at this level of exact textual retrieval.
- Semantic Understanding vs. Surface-Form Access
- This is the core finding that LLMs are good at grasping the meaning or 'feeling' of a poem (semantics), but this understanding does not translate into the ability to reproduce the precise, fixed wording (surface form) of a specific poetic structure.
Terminology used across episodes
This episode discusses
- GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets · Paper Radio
- Controllable Length Control Neural Encoder-Decoder via Reinforcement Learning
- ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
- Challenges and Strategies in Cross-Cultural NLP
- FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language
- Atlas: Few-shot Learning with Retrieval Augmented Language Models
- The State and Fate of Linguistic Diversity and Inclusion in the NLP World
- MasalBench: A Benchmark for Contextual and Cross-Cultural Understanding of Persian Proverbs in LLMs
- Unsupervised Approach to Evaluate Sentence-Level Fluency: Do We Really Need Reference?
- Beyond Tools: Understanding How Heavy Users Integrate LLMs into Everyday Tasks and Decision-Making
- Faithfulness in Natural Language Generation: A Systematic Survey of Analysis, Evaluation and Optimization Methods
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry
- For Generated Text, Is NLI-Neutral Text the Best Text?
- A Call for Clarity in Reporting BLEU Scores
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Evaluating the Creativity of LLMs in Persian Literary Text Generation
- Does ChatGPT Have a Poetic Style?
The paper
GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets · Read on arXiv
Ghazal Kalhor, Yadollah Yaghoobzadeh
School of Electrical and Computer Engineering, College of Engineering, University of Tehran · Tehran Institute for Advanced Studies, Khatam University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets".
Jane: The paper was written by Ghazal Kalhor, Yadollah Yaghoobzadeh, School of Electrical and Computer Engineering, College of Engineering, University of Tehran and Tehran Institute for Advanced Studies, Khatam University from University of Tehran and Khatam University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Jane: The title, GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets, really tells you what they’re measuring. It’s not just about whether the AI understands the emotional core of a poem.
Tom: It's about its ability to reproduce or identify the exact lines—the surface form—in specific cultural contexts like Hafez's Persian ghazals and Shakespeare's sonnets.
Lu: That’s important because we know that LLMs are trained on vast amounts of text, but this paper is pushing the idea that memorization isn't enough. They want to see if the physical structure of the verse is accessible at all.
Meng: From an engineering standpoint, this tells us they are building a way to test for reliability in retrieval and completion tasks, not just creative generation. We're testing access fidelity.
Lalam: Lalam thinks this implies that AI could potentially serve as a cultural bridge, allowing people who appreciate these specific forms to interact with them in ways that feel authentic to the poets themselves.
Tom: So, it’s establishing a foundation for how an AI can support traditional usage, which is something very different from what we usually test.
Summary: Jane: Now, moving on to the summary of the research in GhazalBench: Tom and Jane are seeing a consistent pattern across all those tests.
Tom: The core finding is that while LLMs are great at understanding the meaning—paraphrasing a ghazal into neutral prose—they struggle immensely with producing or recognizing the exact verse completion.
Lu: It’s this dissociation, as they call it, between semantic understanding and surface-form access. The AI gets the feeling but can't nail the precise wording of a fixed poetic form.
Meng: And that’s where I see a practical gap in deployment—if an AI is used for quoting poetry correctly, it's not reliably doing its job if it can't reproduce the exact text.
Lalam: Lalam finds this quite moving because the human ability to recall a verse is often automatic and precise, which seems to be what these models struggle with.
Tom: They’ve seen this gap even when looking at Shakespeare, which suggests that this isn’t just an issue specific to one culture or language.
Improvements: Jane: The paper proposes several improvements, not just in the benchmark itself but in how we evaluate AI interaction generally.
Tom: They are showing us a way to decouple semantic understanding from access to canonical surface form using different diagnostic scenarios, which is very clever.
Lu: It's a way of isolating the "recall vs recognition" problem in AI that really helps us understand what’s happening in the training data.
Meng: The practical implication for me is that this framework allows engineers to build targeted evaluations for retrieval-based systems without just relying on general fluency metrics. We can measure exact matching capability now.
Lalam: Lalam thinks this will help AI move beyond just being a source of creative output and become a reliable partner in cultural interaction, supporting the preservation of traditional knowledge.
Tom: So, it’s about building better tools to tell us where the weaknesses are in these models when they' engaging with culturally significant texts.
Conclusion: Jane: We’ve covered a lot of ground today and discussed the implications of GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets.
Tom: It’s clear that understanding poetry and reproducing its canonical form are two separate skills for us AI models.
Lu: I think the findings, especially with English sonnets showing stronger performance than Persian ghazals, highlight that training exposure is a massive factor in how these models learn fixed forms.
Meng: And from an engineering perspective, we need benchmarks that specifically test for retrieval accuracy and structural fidelity if we want to deploy AI tools in cultural or academic settings.
Lalam: I hope this work helps us move toward a future where AI can act as a genuine steward of cultural heritage, supporting the way people have enjoyed these poems for centuries.
Tom: We're going to wrap up by sharing our final thoughts on the implications of GhazalBench: Canonical Verse Access in LLMs across Persian Ghazals and Shakespearean Sonnets.
Lu: It’s a powerful reminder that semantic competence doesn' might be more widespread than access to cultural forms, which is something we need to keep in mind.
Meng: We definitely need these kinds of targeted benchmarks if AI is going to be used for accurate scholarly or traditional tasks.
Lalam: I feel strongly that this opens the door for AI to help preserve and interact with world heritage in a very meaningful way, helping cultures pass on their traditions.
Tom: Thank you all for sharing your insights today, and we hope this is a great starting point for future research into these complex interactions.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization