LMSpell: Spell Correction with Pre-Trained Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "LMSpell: Spell Correction with Pre-Trained Language Models".
Tom: The gist The first empirical study on Large Language Models (LLMs) for spell correction reveals that LLMs outperform their encoder-based and encoder-decoder counterparts when the fine-tuning dataset is large,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, let's talk a bit about who wrote this, the authors. We’ve got Gunathilakea, Karunarathnaa, Bandaranayakea, and de Silva leading the research here on LMSpell: Spell Correction with Pre-Trained Language Models.
Jane: They are researchers from the Department of Computer Science and Engineering at the University of Moratuwa in Sri Lanka and also from Massey University over in New Zealand. It’s a collaboration across different institutions, which is pretty common when you're tackling these complex language problems.
Lu: The team has a solid background in computational science and language modeling, which gives them the foundation to do this kind of comprehensive comparison across different model architectures.
Meng: From an engineering standpoint, seeing researchers from different places working together on a generalized toolkit like LMSpell is important because it shows the need for tools that aren't tied to just one specific research lab.
Lalam: It points toward a future where spell correction tools become much more flexible and less dependent on the specific pre-training of the underlying language model.
The paper's summary: Tom: Okay, so what is the main idea here in LMSpell: Spell Correction with Pre-Trained Language Models? It basically presents an empirical study to see how effective pre-trained language models are for spell correction, specifically looking at low-resource languages.
Jane: The core finding they highlight is that when you fine-tune these models with a large dataset, the large language models beat both the encoder-based and encoder-decoder versions of those same models.
Lu: They're doing this by testing different model types—encoder-based, decoder-based, and encoder-decoder—across a wide range of languages to see where each architecture shines or struggles.
Meng: It sounds like they are showing that the way you structure the model actually matters depending on what kind of data you feed it during fine-tuning.
Lalam: And they also introduce this new script for spell correction that has a built-in mechanism to handle hallucinations, which is a big deal for reliability.
The paper's improvements: Tom: Now let's look at the specific improvements they suggest with LMSpell. One of the biggest things is introducing this new evaluation script that deals with hallucination issues in spell correction.
Jane: They noticed that previous research didn't properly align sentences when a model makes a mistake and it just messed up the whole sentence structure, which really distorted metrics like precision and recall.
Lu: So this new script aligns the original, predicted, and expected sentences at the word level first to spot those words that are hallucinated in the prediction, treating them as false positives.
Meng: And then for everything else left over, they align on a character level to compare them, which keeps the standard metrics much more honest about what's actually working.
Lalam: This capability is crucial because it makes the evaluation of these spell correctors way more robust and less prone to being tricked by structural errors introduced by the AI.
Conclusion: Tom: So, wrapping up LMSpell: Spell Correction with Pre-Trained Language Models, the main thing we see is that for low-resource languages not included in an LLM's initial training, large language models can achieve better results if you fine-tune them with a substantial task-specific corpus.
Jane: It’s not just about which model architecture you use; it’s about having enough data to train that model effectively for the specific language task at hand.
Lu: The case study on Sinhala spell correction they did really opens up possibilities for how we can improve spelling accuracy in languages that haven't been well-covered by current models.
Meng: I think from a practical standpoint, this suggests that if we have the right training data, leveraging these larger language models for specialized tasks like spell checking is a viable path forward.
Lalam: And the LMSpell toolkit itself gives us an easy way to implement these findings across all PLMs, which makes this research immediately usable for researchers worldwide.
Tom: So that’s our take on LMSpell: Spell Correction with Pre-Trained Language Models, showing that big models can be powerful tools for correcting spelling in many different languages when you give them the right training material.
Department of Computer Science and Engineering, University of Moratuwa · School of Mathematical and Computational Sciences, Massey University
cs.CL
Submitted: 2025-12-05
Updated: 2026-10-08
Code: https://github.com/huggingface/accelerate
Importance score: 80/100
The gist: The gist The first empirical study on Large Language Models (LLMs) for spell correction reveals that LLMs outperform their encoder-based and encoder-decoder counterparts when the fine-tuning dataset
Key concepts
- Encoder-based Model
- These models treat spell correction as a sequence labeling task using masked language modeling. They are adapted by assigning each input token an output token: the correct word is labeled as itself, while an incorrect word is labeled with its properly corrected form.
- Decoder-based Model
- This approach uses modern LLMs where the model generates an output based on a prompt. The prompt instructs the LLM to act as an expert spell corrector, prompting it to generate the corrected text in response to a given sentence.
- Hallucination Handling
- To prevent incorrect content insertion (hallucination), a new evaluation script aligns sentences at the word level first. This identifies predicted words that are not in the original or expected sentences as false positives. Then, character-level alignment is used for comparison, ensuring metrics reflect true structural accuracy.
Terminology
Summary
The gist The first empirical study on Large Language Models (LLMs) for spell correction reveals that LLMs outperform their encoder-based and encoder-decoder counterparts when the fine-tuning dataset is large, even in languages where the LLM was not pre-trained.
Challenges for Low-Resource Languages
Spell correction remains a challenging problem for low-resource languages (LRLs) because pretrained language models (PLMs) have been limited to a handful of languages, and there has been no proper comparison across PLMs. For high-resource languages (HRLs), spell correction is more or less a solved problem. For the LRL Sinhala, there exists only a handful of research on spell correction, and their accuracy is not par for practical use.
LMSpell Toolkit
LMSpell is presented as an easy-to-use spell correction toolkit across PLMs that can be extended for any language and any type of PLM. It supports the implementation of spell correctors using all three types of PLMs (encoderbased, decoder-based, and encoder-decoder). The toolkit abstracts basic PLM-related functions such as finetuning and inference, allowing users to select a model without needing to consider model-specific implementation details. The pipeline integrates Accelerate (Gugger et al., 2022) with DeepSpeed (Rasley et al., 8) for efficient fine-tuning of encoder-based and encoder-decoder-based models.
Model Architectures
LMSpell supports three types of PLMs:
-
Encoder-based: These PLMs are adapted by framing spell correction as a sequence labelling task. To implement both error detection and correction only using an encoder-based model, the system follows Hong et al. (2019) and implements sequence labeling such that each input token is assigned a corresponding output token. If the token is correct, it is labeled with itself, whereas an erroneous token is labeled with its corrected form.
-
Decoder-based: Modern LLMs produce an output in response to a textual prompt. The prompt structure involves instructing the model as an expert spell corrector on a given sentence.
-
Encoder-decoder-based: The PLM is fine-tuned to generate the spell-corrected sentence when an erroneous sentence is given as input.
Handling Hallucination
Hallucination in this context refers to cases where a PLM introduces content that is neither in the original nor the expected sentences. These hallucinations shift the word alignment of the predicted sentence compared to the original and expected sentences. To address this, a new evaluation script was developed that first aligns sentences at the word level to identify any hallucinated words in the predicted sentence and treats them as false positives. After flagging those words, the remaining words are aligned on a character level to compare with each other. This ensures that standard metrics such as precision, recall, and F1-scores are not significantly distorted by insertions and deletions that alter the sentence structure.
Empirical Evaluation
The experiments experimented with various PLMs including encoder-based (XLM-R), encoder-decoder (mT5, mBART50), and decoder-based (Llama 3.1 8B, Gemma 2 9B). The models were fine-tuned with 5k sentences for each language. Results showed that even if a language is not included in an LLM, better results than those of encoder or encoder-decoder models can be achieved if the language script is Latin or a derivative thereof. Furthermore, in all languages, including those for which such genesis from Latin script cannot be claimed, LLMs fine-tuned with a large task-specific corpus outperform the other models. The case study on Sinhala spell correction demonstrated possible improvements to spell correction with LLMs.
Performance Across Languages
Table 12 presents performance across different dataset sizes for Sinhala and Hindi, showing that Gemma 2 shows the best result for Sinhala when using the full dataset. For Hindi, Gemma 2, Llama 3.1, and mBART50 fine-tuned with the full dataset show near-equal results. The results suggest that for languages with Latin (or a derivative of it) script (Azerbaijani, Bulgarian, French and Vietnamese), LLMs show remarkable generalization capabilities. However, they struggle with other scripts such as Devanagari (Hindi) and Sinhala. Mixed language training significantly boosts the results of such languages.
Conclusion
LMSpell provides a spell corrector library based on PLMs, and empirical evaluation confirms that for LRLs unrepresented in LLMs, LLMs outperform other types of PLMs when the fine-tuning dataset is large. The experiments with Sinhala as a case study highlight the challenges and benefits of spell correction for LRLs. Future plans include experimenting with further enhancements such as preference optimization.
The gist LMSpell is an easyto-use spell correction toolkit across PLMs that includes an evaluation function that compensates for the hallucination of LLMs.
How it works
-
LMSpell supports the implementation of spell correctors using all three types of PLMs (encoderbased, decoder-based, and encoder-decoder).
-
Encoder-based models are adapted by framing spell correction as a sequence labelling task. This is achieved by labeling each input token with itself if correct and its corrected form if erroneous using masked language modeling (MLM).
-
Decoder-based models are prompted with instructions such as You are an expert language spell corrector to generate the corrected output.
-
Encoder-decoder-based PLMs are fine-tuned to generate the spell-corrected sentence when given an erroneous input sentence.
Evaluation Workflow
The evaluation workflow during hallucination involves aligning the original, predicted, and expected sentences at the word level to identify any hallucinated words in the predicted sentence and treating them as false positives. After flagging these words, the remaining words are aligned on a character level to compare with each other. This methodology avoids the issue where previous research lacked proper string alignment and makes standard metrics like precision, recall and F1-scores more representative of the model's true performance.
Training Parameters
Training parameters for Encoder-based and Encoder-Decoder models are detailed in Table 9. For LLMs, training parameters are detailed in Table 10. The experiments utilized different GPU types, including NVIDIA T4 and L4 GPUs.
Language Dataset Details
Table 11 shows the details of the languages used in this study, categorizing them by resource level according to Ranathunga and de Silva (2022). The resource level is categorized from Category 0 to Category 5. The dataset sizes for Sinhala and Hindi are detailed in Table 12.
Performance Metrics
Performance is reported using F0.5-score for the correction task and F1-score for the detection task. Table 7 shows performance for different error percentages of the test set, while Table 12 shows performance across different dataset sizes for Sinhala and Hindi.
Limitations
Due to the limitations in computing resources, experiments considered only 5 PLMs and only 7 spell correction datasets. Dataset size-based experiments were limited to only Sinhala and Hindi due to computer resource limitations. The biases in existing datasets and PLMs may reflect on the spell corrector output.
Improvements for AI systems
-
The LMSpell toolkit can be extended for
any language and any type of PLM,
enabling researchers to apply spell correction methods beyond those currently pre-trained in existing frameworks like NeuSpell. This allows for rapid prototyping of new spell correction strategies across diverse linguistic contexts, including low-resource languages (LRLs) not yet covered by major LLMs. -
The new script introduced within LMSpell
takes hallucination into consideration,
which is critical becausehallucination refers to cases where a PLM introduces content that is neither in the original nor the expected sentences.
This capability allows the system to specifically flag insertions or deletions that alter sentence structure, ensuring that metrics like precision and recall are notsignificantly distorted
by model-introduced errors. -
LLMs, when fine-tuned with a large dataset,
outperform other types of PLMs,
especially for LRLs where they arenot pre-trained with Si data.
This suggests that LLM architectures (like Gemma 2) are highly effective for LRL spell correction when provided with sufficient task-specific corpora, potentially leading to more robust and generalizable correction than encoder or encoder-decoder models. -
For languages like Sinhala, the system can leverage
in-context learning and RAG
techniques, as shown in Table 3 results whereRAG 1
andRAG 2
yielded high scores. This allows the AI to generate contextually relevant correction candidates or similar sentences from a vector store during inference, improving performance where traditional fine-tuning alone might fall short. -
The evaluation methodology moves beyond naive character comparison by aligning sentences
at the word level, identifying any hallucinated words in the predicted sentence and treating them as false positives.
This ensures that standard metrics are not compromised by positional shifts, which previously causeda cascading effect of mismatches across the remainder of the sequence.
Sources
- A Survey on Transfer Learning in Natural Language Processing
- SinLlama -- A Large Language Model for Sinhala
- Survey on Publicly Available Sinhala Natural Language Processing Tools and Research
- The Faiss library
- Data Augmentation and Terminology Integration for Domain-Specific Sinhala-English-Tamil Statistical Machine Translation
- On the (In)Effectiveness of Large Language Models for Chinese Text Correction
- A Comprehensive Approach to Misspelling Correction with BERT and Levenshtein Distance
- KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
- Sentiment Analysis for Sinhala Language using Deep Learning Techniques
- Multilingual Translation with Extensible Multilingual Pretraining and Finetuning
- Gemma 2: Improving Open Language Models at a Practical Size
- Does Correction Remain A Problem For Large Language Models?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering