Survey on Publicly Available Sinhala Natural Language Processing Tools and Research
cs.CL
Submitted: 2019-06-05
Updated: 2026-09-06
Code: https://github.com/lknlp/lknlp.github.io
Project page: https://openslr.org/52
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Geographically-Informed Language Identification
- The FLoRes Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-English
- Sinhala Language Corpora and Stopwords from a Decade of Sri Lankan Facebook
- Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
- Exploiting Parallel Corpora to Improve Multilingual Embedding based Document and Sentence Alignment
- Google Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An Overview
- NSINA: A News Corpus for Sinhala
- Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages
- Adapting the Tesseract Open-Source OCR Engine for Tamil and Sinhala Legacy Fonts and Creating a Parallel Corpus for Tamil-Sinhala-English
- FastText.zip: Compressing text classification models
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Stylomech: Unveiling Authorship via Computational Stylometry in English and Romanized Sinhala
- SOLD: Sinhala Offensive Language Dataset
- Predicting the Type and Target of Offensive Posts in Social Media
- XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented Languages
- An Open Dataset and Model for Language Identification
- Sinhala-English Word Embedding Alignment: Introducing Datasets and Benchmark for a Low Resource Language
- CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
- MADLAD-400: A Multilingual And Document-Level Large Audited Dataset
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering