MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
cs.CL
Submitted: 2025-10-09
Updated: 2025-10-09
Comments: Pre-print
Journal ref: https://aclanthology.org/2026.wildre-1.1/
Project page: https://chandamu.github.io/https://sanskrit.iitk.ac.in/jnanasangraha/chanda
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Prior NLP work studying poetry has focused primarily on automatic poem generation and summarization.
Terminology
Abstract
Prior NLP work studying poetry has focused primarily on automatic poem generation and summarization. Many languages have well-studied traditions of poetic meter which enforce constraints on a poem in terms of syllable and phoneme patterns. Such advanced literary forms offer opportunities for probing deeper reasoning and language understanding in Large Language Models (LLMs) and their ability to follow strict pre-requisites and rules. In this paper, we introduce MetricalARGS, the first taxonomy of poetry-related NLP tasks designed to evaluate LLMs on metrical poetry across four dimensions: Analysis, Retrieval, Generation, and Support. We discuss how these tasks relate to existing NLP tasks, addressing questions around datasets and evaluation metrics. Taking Telugu as our example language, we illustrate how the taxonomy can be used in practice. MetricalARGS highlights the broader possibilities for understanding the capabilities and limitations of today's LLMs through the lens of metrical poetry.
Sources
- Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs
- Chandomitra: Towards Generating Structured Sanskrit Poetry from Natural Language Inputs
- PoeTone: A Framework for Constrained Generation of Structured Chinese Songci with LLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering