Data Synthesis and Parameter-Efficient Fine-Tuning for Low-Resource NMT: A Case Study on Q'eqchi' Mayan
cs.CL, cs.AI, cs.LG
Submitted: 2026-06-08
Updated: 2026-06-08
Comments: Accepted to the 29th International Conference on Text, Speech and Dialogue (TSD 2026). This version of the contribution has been accepted for publication, after peer review but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections
Journal ref: Text, Speech, and Dialogue (TSD 2026), Lecture Notes in Computer Science, vol. 16940, pp. 188-200, Springer, 2027
DOI: 10.1007/978-3-032-37249-9_16
Code: https://github.com/achulzhanov/mayan-mt5
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- mT5: A massively multilingual pre-trained text-to-text transformer
- R2T: Rule-Encoded Loss Functions for Sequence Tagging in Low-Resource Languages
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering