SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0".
Jane: SiDiaC-v.2.0 represents the largest comprehensive Sinhala diachronic corpus to date,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So the team is absolutely buzzing about this paper on SiDiaC-v.two point zero: Sinhala Diachronic Corpus Version two point zero, which they’ve put out on arXiv and it's huge for our field right now. Jane, can you give us the quick rundown of what this whole corpus actually is?
Jane: Absolutely, Tom. Basically, the main thing we need to know is that SiDiaC-v.two point zero represents the largest comprehensive Sinhala diachronic corpus to date and it spans a massive historical range from one thousand eight hundred CE to one thousand nine hundred fifty-five CE in terms of publication dates, covering a historical span from the 5th to the 20th century CE in terms of written dates. It consists of twenty-four hundred and forty-one thousand words spread across one hundred eighty-five literary works that went through serious filtering, preprocessing, and copyright compliance checks before it was finalized.
Lu: That scale is what really gets my creative gears turning; having such a massive dataset allows for so much deeper structural analysis than we've ever managed in this area. I'm particularly interested in how the annotation layers are structured because that’s where the real linguistic potential lies.
Meng: From an engineering standpoint, that level of data volume is impressive, but I wonder about the practical implications for our current processing pipelines; how much computational overhead are we looking at when handling twenty-four hundred and forty-one thousand words across one hundred eighty-five literary works?
Lalam: I think from an AI perspective, the sheer breadth of historical and linguistic data makes it incredibly valuable for training models to understand nuanced cultural contexts, which is where I see the biggest potential for improving our understanding of Sinhala culture.
Tom: That’s a great point, Lalam; understanding those cultural contexts is exactly what we hope this resource enables us to do better than before. Jane, following up on that massive scope, why does the historical span from the 5th to the 20th century CE matter so much for our research?
Jane: That historical depth allows researchers to trace linguistic changes across centuries, giving us a view of how language evolves over a very long period, which is essential for diachronic studies. It’s not just about the words themselves; it’s about seeing the language in its full historical context.
Lu: And looking at how this corpus compares to other resources, like FarPaHC or COHA, we see a significant gap in scale here, which makes SiDiaC-v.two point zero a unique benchmark for linguistic resource size. I can imagine the architectural challenges of integrating such a vast and varied set of texts into our existing models.
Paper summary: Meng: Architecturally, I'm thinking about the preprocessing steps; they mentioned using Google Document AI OCR initially, which suggests we have to deal with real-world noise like formatting errors and code-mixing before the actual linguistic analysis can even start. How robust is that initial digitisation process for historical Sinhala text?
Lalam: The post-processing steps mentioned, which include correcting spacing errors, fixing multi-column text, and addressing code-mixing, are crucial because those kinds of artifacts can severely skew automated analysis if not handled carefully. These corrections help ensure the subsequent linguistic tagging is as accurate as possible for cultural preservation.
Tom: It sounds like the authors put a lot of work into cleaning up the raw data to make sure we aren't just feeding noise into our systems, which is smart thinking. But what about how they handled the written dates, since that’s so critical for a diachronic corpus?
Jane: That’s where they made some key revisions; initially, their criteria for written date annotation were too strict and relied heavily on one source, but they adjusted the methodology to adopt an approach similar to COHA, which includes using author lifespans when possible. This adjustment allowed them to include a broader range of literary works in this version.
Lu: That shift in date annotation strategy is significant because it broadens the scope of works we can use for historical comparison, moving beyond just rigidly defined sources. It opens up new avenues for comparing linguistic features across different chronological periods within the corpus.
Meng: From a practical standpoint, if we integrate this into our current tools, I need to know how much manual intervention is still required after all that filtering and post-processing is done so we can actually get usable outputs quickly. The efficiency of the annotation layers will determine our adoption rate here.
Lalam: I see the potential for this corpus to significantly enhance how we model cultural evolution; imagine an AI that can track specific linguistic features through centuries of religious or poetic texts, which could provide incredible insights into the history of Theravada Buddhism in Sri Lanka.
Tom: That’s exactly what I was thinking—seeing those deep cultural patterns emerge through data analysis. Now, let's move toward the conclusions section and what this means for the broader world outside our immediate research area. Jane, what’s your take on the overall impact of SiDiaC-v.two point zero?
Jane: When we look at the title and authors of this work, Nevidu Jayatilleke, Nisansa de Silva, Uthpala Nimanthi, Gagani Kulathilaka, Azra Safrullah, and Johan Sofalas from the Department of Computer Science and Engineering at the University of Moratuwa do show a strong academic foundation for this kind of large-scale linguistic project. The implications are that we now have a far more robust tool to study Sinhala language history and development over a long timeline.
Paper summary: Lu: I think the most significant implication is that this corpus provides a high-quality, large-scale foundation for building sophisticated AI models specifically tuned for low-resource languages like Sinhala, which pushes the boundaries of what we can achieve in computational linguistics. It gives us a substantial resource to test and refine new parsing and generation techniques.
Meng: I’m focusing on the practical application here; if we can successfully deploy models trained on this data, it could mean that future language processing tools for Sinhala will be far more accurate in real-world applications than they are today. That accessibility to high-quality data is what matters most to me right now.
Lalam: From a cultural standpoint, the impact of SiDiaC-v.two point zero could be immense because it gives us the means to preserve and analyze Sinhala literary traditions with unprecedented detail, which is vital for understanding our unique cultural heritage. This data repository is a treasure for future generations of scholars and creators alike.
Tom: So, to wrap up on the conclusion of this discussion about SiDiaC-v.two point zero, we’ve seen that this corpus is massive, spanning centuries and covering over one hundred eighty-five literary works. The authors have done a lot of work to ensure the data quality through rigorous filtering and post-processing, which includes fixing things like spacing errors and addressing code-mixing.
Jane: And it’s important to remember that the methodology itself was refined, particularly in how they handled written dates by incorporating author lifespans alongside other sources, which makes the temporal analysis much richer. This work provides a substantial resource for anyone interested in Sinhala diachronic studies.
Lu: The future work suggested by the paper points toward using this corpus to test new techniques for language modeling and semantic understanding, which means we’ve got a very fertile ground for pushing those computational limits forward. We can really start building on top of this foundation.
Meng: And I just want to reiterate that the practical impact hinges on how smoothly these large-scale datasets can be integrated into existing infrastructure, so we need to keep thinking about those computational realities as we move forward.
Lalam: Ultimately, SiDiaC-v.two point zero gives us the raw material to build an AI that can deeply understand and reflect the nuances of Sinhala culture through its historical literature, which is a powerful thing to have. This is about empowering language understanding for our community.
Conclusion: Tom: So we've been deep into the details of SiDiaC-v.two point zero, and now it's time to look at what this whole thing actually means for us in the broader academic world.
Jane: That’s right, Tom; we’re going to wrap up by discussing the title and those authors and what this corpus really implies for future research directions.
Lu: I think focusing on the scope of one hundred eighty-five literary works spanning such a long period opens up completely new ways to visualize linguistic evolution across Sri Lanka.
Meng: From my side, I'm thinking about how this massive dataset will affect the practical deployment of AI tools for Sinhala, specifically in terms of accuracy and handling historical text variations.
Lalam: I see the potential here as a way to build an AI that can genuinely understand the deep cultural layers embedded in our literature, which is something truly significant for preserving heritage.
Tom: Exactly, Lalam; this isn't just about counting words, it's about building a much richer understanding of how Sinhala culture has been expressed over centuries.
Jane: The authors themselves are from the Department of Computer Science and Engineering at the University of Moratuwa, and their academic background really shows why they tackled such a complex project.
Lu: Their expertise in computer science is clearly what made them capable of handling the immense data processing and annotation layers we saw earlier.
Meng: That technical skill is what allows them to navigate those OCR errors and formatting issues we talked about, making the final product usable by engineers like us.
Lalam: And for me, it’s the way this corpus gives us a foundation to train models that don't just translate words but actually grasp the cultural context behind them.
Tom: That’s a powerful vision, Lalam; we’re looking at tools that can capture history in language itself, and that kind of capability is what keeps me energized about this work.
Jane: It really shows how meticulous the groundwork has been done to ensure this corpus is reliable enough for serious diachronic studies across the entire historical spectrum.
Lu: We should definitely keep an eye on their future plans mentioned in the paper, because they pointed toward using this data to test new language modeling techniques.
Meng: I agree; testing those new models against this high-quality historical data will give us concrete benchmarks for performance in real-world Sinhala applications.
Department of Computer Science & Engineering, University of Moratuwa, Sri Lanka · Research Department, Informatics Institute of Technology, Sri Lanka
cs.CL
Submitted: 2026-03-11
Updated: 2026-10-02
Code: https://github.com/NeviduJ/SiDiaC-v.2.0
Importance score: 67/100
The gist: SiDiaC-v.2.0 represents the largest comprehensive Sinhala diachronic corpus to date, spanning from 1800 CE to 1955 CE and covering a historical span from the 5th to the 20th century CE, making it a
Key concepts
- Diachronic Corpus
- A large collection of texts spanning a long period in time. This corpus covers 1800 to 1955 CE, allowing researchers to study how the language and literature have changed over centuries in Sinhala.
- Corpus Construction
- The process of building the SiDiaC-v.2.0 dataset involved several steps: collecting data from sources like the National Library, using Google Document AI OCR for digitization, and applying rigorous filtering to ensure quality and relevance.
- Genre Analysis
- This involves categorizing texts into types like Poetry or Religious works. The analysis shows that religious and poetry genres are most prevalent, reflecting the deep cultural link between Sinhala literature and Theravada Buddhism.
Terminology
Summary
SiDiaC-v.2.0 represents the largest comprehensive Sinhala diachronic corpus to date, spanning from 1800 CE to 1955 CE and covering a historical span from the 5th to the 20th century CE, making it a vital resource for Sinhala Natural Language Processing and temporal linguistic studies.
Corpus Construction and Scale
The SiDiaC-v.2.0 corpus is characterized by its massive scale, consisting of 241k words across 185 literary works.
The construction process involved several meticulous steps to ensure quality, building upon the foundation of SiDiaC-v.1.0. Key stages included:
(i) Data Collection and Sourcing:
The corpus was informed by practices from other corpora such as FarPaHC, SiDiaC-v.1.0, and CCOHA to guide strategies for syntactic annotation and text normalization in a low-resource language context. Texts were primarily sourced from the National Library of Sri Lanka (SiDiaC-v.1.0) and expanded through careful filtration mechanisms to reach 185 literary works in SiDiaC-v.2.0, moving beyond the initial 46 works of the previous version due to revised filtration criteria that adopted a similar approach to COHA for written date annotation.
(ii) Preprocessing and Digitization:
The initial extraction utilized Google Document AI OCR, which was selected for its superior real-world accuracy, especially concerning text modernisation and morpheme segmentation for historical Sinhala.
Extensive post-processing was then conducted to correct formatting issues such as spacing errors, multi-column text, misplaced words and phrases,
and to address code-mixing.
(iii) Annotation Layers:
The corpus is categorized into two layers: a primary binary categorization (Non-Fiction or Fiction), and a secondary detailed categorization grouping texts under specific genres such as Religious, History, Poetry, Language, and Medical.
Metadata creation includes fields like the title in Sinhala/Romanized form, author names (with romanized versions for unknown authors), genre classifications at both levels, issued date, written date (annotated based on author lifespans where applicable), and OCR confidence scores.
Data Filtration and Quality Control
Significant effort was dedicated to refining the dataset through rigorous filtration pipelines to address limitations identified in SiDiaC-v.1.0:
(i) Written-Date Annotation Revision:
The filtration criteria were revised because the initial strict criteria in SiDiaC-v.1.0 led to an overreliance on a single source for written dates (relying heavily on Sannasgala, 2015). The methodology was adjusted to adopt a similar approach to COHA, including author lifespans when applicable, enabling the inclusion of a broader range of literary works in SiDiaC-v.2.0.
(ii) Text and Content Filtering:
Documents were filtered based on copyright laws (Intellectual Property Act No. 36 of 2003), focusing on literature by authors who passed away before 1955 or works by unknown authors published before that year, which resulted in the removal of 20 documents
from the initial set. Furthermore, documents containing entirely non-Sinhala text and those that contained only tabular information
were removed, resulting in the final set of 185 books.
(iii) OCR Error Correction:
The post-processing step addressed character-level errors identified during OCR, such as spell correction domain errors,
where substitutions or deletions occurred due to noise in scanned documents. The process involved examining the entire corpus character by character and addressing formatting issues simultaneously with these corrections for efficiency.
Genre Analysis and Temporal Distribution
The genre analysis of SiDiaC-v.2.0 reveals a predominance of certain genres, reflecting cultural influences:
(i) Genre Prevalence:
The analysis based on publication dates shows a right skew in the publication counts,
with the highest number of publications occurring in the period from 1880 to 1900, including 63 books across various secondary genres. Overall, 141 of 185 books belong to either Poetry (54) or Religious (86) genres.
(ii) Century-wise Distribution:
When analyzing documents by written century and genre, the 20th century contained the largest number of books, totaling 17. The analysis shows that religious texts and poetry being the most prevalent
stems from the close relationship between Sinhala literary culture and Theravada Buddhism,
alongside the influence of Sanskrit Kavya traditions.
Improvements for AI systems
Here are specific improvements for AI systems derived from the SiDiaC-v.2.0 corpus:
-
A low-resource language (Sinhala) Named Entity Recognition (NER) model trained on 67,005 words of written-date annotated text can be developed.
-
The improved NER system can accurately identify and classify entities related to historical figures, religious texts, and literary genres within Sinhala literature with high precision.
-
A diachronic word embedding model for Sinhala can be trained using the 80 consistent words identified from the 13th to 20th centuries (e.g.,
සතර
, "මහ"). -
The improved word embedding system will allow researchers to track semantic shifts of key historical and religious concepts (like 'wisdom' or 'greatness') over time in Sinhala texts, providing a temporal linguistic analysis tool.
-
A text normalization pipeline can be optimized using the manual post-processing rules derived from SiDiaC-v.2.0 (e.g., handling code-mixing between Pali/Sanskrit and Sinhala) to generate cleaner, more reliable input for downstream NLP tasks like machine translation or semantic parsing in low-resource settings.
-
A
Zero-Shot OCR Accuracy
assessment tool can be built by comparing the performance of Document AI versus Surya on historical Sinhala texts, specifically focusing on text modernization and morpheme segmentation capabilities, to determine optimal pre-processing workflows for archival digitization. -
A Genre Classification Model (Primary/Secondary) can be fine-tuned using the 185 categorized works to predict genre automatically, enabling large-scale corpus mining for specific literary themes (e.g., predicting 'Religious' or 'Poetry' content).
-
The system can perform
Temporal Text Provenance
analysis by cross-referencing the publication dates and written dates of documents to determine the temporal relationship between an original text and its commentary, mitigating ambiguity in historical source attribution. -
A sophisticated filtering mechanism for data curation can be implemented, mirroring the successful 4.3 filtration steps (copyright, non-Sinhala text removal), to automatically curate large, noisy digital archives into high-quality analytical datasets suitable for NLP model training.
Sources
- Survey on Publicly Available Sinhala Natural Language Processing Tools and Research
- Sinhala Language Corpora and Stopwords from a Decade of Sri Lankan Facebook
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering