Padamitra: Grounded Glossary Generation for Classical Sanskrit

arXiv:2608.25038 · cs.CL · Submitted 2026-08-25 · Read on arXiv

cs.CL

Submitted: 2026-08-25

Updated: 2026-09-01

Comments: Accepted in the Findings of EMNLP 2026

Code: https://github.com/sanganaka-iitkgp/Padamitra-Glossary-Generation

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation

Terminology

Abstract

We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective. We construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency. Across zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outperforms prompting, while explicit segmentation yields gains. Error analysis identifies over-segmentation of sandhi and samasa compounds as the dominant failure mode, pointing to morphological modeling as the key bottleneck for faithful Sanskrit lexical decomposition.

Sources

Related papers