Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text

arXiv:2608.18437 · cs.CL · Submitted 2026-08-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Tangut Word Segmentation under Extreme Resource Scarcity".

Tom: Tangut word segmentation under extreme resource scarcity is addressed by integrating traditional lexicons and unlabeled text within a unified BIES–CRF framework.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, diving into the specifics of "Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text," the authors are essentially saying they've developed a first systematic study for automatically segmenting Tangut words using expert annotations, traditional lexicons, and unlabeled text.

Jane: That’s right, Tom. The core idea is that Tangut script doesn't mark word boundaries explicitly at all, which makes segmentation incredibly hard without much prior knowledge. They are presenting this as the first organized effort to solve that problem by using a combination of these three distinct knowledge types.

Lu: The paper highlights their contribution in three main areas: first, they define the task and evaluation method, which is using character-level BIES sequence labeling to find word boundaries; second, they propose a way to integrate lexicon features that are reliability-calibrated; and third, they systematically compare different methods for learning from unlabeled Tangut text.

Meng: I'm looking at the data resources mentioned—they used two thousand seven hundred fifty expert-annotated segments and thirty-one thousand eight hundred ninety-three word tokens from texts like Buddhist scriptures and the secular encyclopaedic work Leilin. The fact that secular documents make up about ninety-one point five percent of all segments is a significant detail for anyone thinking about practical application.

Lalam: That high proportion of secular text tells us that most of the work, and likely the practical usage, will be in those more common documents, which gives us a good starting point for testing how well the system generalizes across different textual styles.

The paper's summary: Tom: So, summarizing what they found in this "Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text," the researchers demonstrated that combining expert annotation with their lexical-lattice features and contextual masked language modeling pretraining gives the best results.

Jane: That's because their final setup, which includes the complete TangutEncoder, achieves a mean F1 of zero point nine one one under within-source line-level evaluation, which is quite strong given the constraints they were working with. This score shows that adding those contextual features really helps boost performance beyond what the basic lexicon and statistical features provide alone.

Lu: What's striking is that the full TangutEncoder manages to learn contextual character representations without needing to assume word boundaries beforehand, which is a key technical feat for this kind of language where boundaries are implicit. It seems they successfully trained the encoder using single-character and contiguous-span masking, selecting about fifteen percent of positions with mixed masking types.

Meng: I'm interested in the comparison they made between their different knowledge components; they showed that while the linear CRF benefited from lexical and statistical features, it was the full TangutEncoder that actually attained the highest mean score and OOV recall across thematically diverse passages.

Lalam: It really shows that you get a better result when you integrate all those pieces—the lexicon, the statistics, and especially that contextual pretraining—rather than relying on just one part of the system. That holistic approach is what drives the superior performance they reported in their study.

The paper's improvements: Tom: Moving into what the paper suggests for future improvements, they are focused on refining how those different knowledge sources interact, specifically by proposing a reliability-calibrated lexicon-lattice representation.

Jane: That means they aren't just using a standard dictionary; they are giving each dictionary candidate a score based on how reliable it is through out-of-fold estimation, which prevents the system from getting biased by its own training data statistics for every single instance.

Lu: They also suggest incorporating explicit distributional features directly into the supervised layer, like log bigram frequency and character association strength, to give the sequence labeling model more immediate context from the unlabeled text before it even gets through the contextual encoder.

Meng: From an implementation standpoint, adding these real-valued feature functions into the linear CRF layer is straightforward for a standard CRF setup; it just means feeding those calculated features as extra inputs during decoding. I wonder if that complexity translates to usable performance gains in a live deployment environment where you need low latency.

Lalam: The suggestion to use an aggregated lexicon-lattice representation sounds like it helps manage the sheer volume of overlapping dictionary candidates effectively, which is crucial when dealing with a script where many characters can combine in various ways.

Conclusion: Tom: So, wrapping up the findings from this "Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text," we see that the full system, TangutEncoder plus the dictionary and distributional features, achieved a mean F1 of zero point nine one seven±zero point zero zero three in secular text, while the CRF performed stronger on religious texts.

Jane: That distinction is important because it shows that context learning through pretraining really helps improve performance across different genres when you train the entire encoder rather than just using static character embeddings, which is a key point they made about their findings.

Lu: The paper emphasizes that the contextual pretraining substantially improves F1 and OOV recall for both religious and secular texts by training the entire encoder with MLM, clarifying that this comprehensive approach is what drives those gains over simpler methods.

Meng: For practical application, it suggests that to get the best results in a deployment scenario, you need to consider whether you prioritize maximizing performance on one domain or balancing performance across both types of text distributions.

Lalam: The implication for the future is that this methodology provides a solid blueprint for tackling segmentation problems in other low-resource or script-less languages by systematically testing how lexical, statistical, and contextual knowledge can be fused effectively.

Tom: Exactly. We've seen how they used expert annotation to build the foundation and then layered on sophisticated AI techniques to solve a real linguistic challenge. That's what we have today with this work on Tangut word segmentation under extreme resource scarcity.

Key Laboratory of Linguistics, Chinese Academy of Social Sciences (University of Chinese Academy of Social Sciences) · Institute of Linguistics, Chinese Academy of Social Sciences · Institute of Ethnology and Anthropology, Chinese Academy of Social Sciences · School of Software and Microelectronics, Peking University

cs.CL

Submitted: 2026-08-19

Updated: 2026-10-01

Code: https://github.com/jiangli-va/TangutSeg

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: Tangut word segmentation under extreme resource scarcity is addressed by integrating traditional lexicons and unlabeled text within a unified BIES–CRF framework.

Key concepts

BIES Sequence Labeling
This is a task where a model tags every character in a text to determine if it starts (B), is inside (I), or ends (E) a word, or if it's a single-character word (S). The goal is to predict the correct sequence of these tags for an entire line of text.
TangutEncoder
This is a specialized BERT-style encoder designed specifically for the Tangut writing system. It uses character-level masked language modeling, where parts of characters are hidden and the model tries to guess them based on their surrounding context, learning rich contextual representations without knowing word boundaries beforehand.
Lexicon-Lattice Features
This layer uses a structured representation of known dictionary words to provide features for the segmentation model. It encodes overlapping dictionary candidates into binary features that help the system decide which character positions are likely part of a valid word based on their presence in the lexicon.

Terminology

Summary

Tangut word segmentation under extreme resource scarcity is addressed by integrating traditional lexicons and unlabeled text within a unified BIES–CRF framework. The core finding demonstrates that combining expert annotation, lexicon-lattice features, and contextual masked language modeling pretraining yields the highest performance for automatic word segmentation in Tangut, achieving a mean F1 of 0.911 under within-source line-level evaluation.

The gist: The full TangutEncoder reaches the highest mean F1 (0.911) and improves recall beyond the labeled training vocabulary, demonstrating generalization beyond the limited supervised vocabulary across thematically diverse held-out passages.

Task Formulation and Data Resources

The task is formulated as character-level BIES sequence labeling, where B, I, and E denote the beginning, inside, and end of a multi-character word; S denotes a single-character word. The model predicts the highest-scoring valid sequence: ˆy = arg max y∈Y(x) pθ(y x). The data resources consist of an expert-annotated corpus containing 2,750 textual segments and 31,893 word tokens from a Buddhist scripture and the secular encyclopaedic work Leilin. The genres differ substantially in size and lexical distribution, with secular documents accounting for approximately 91.5% of all segments.

Layered Framework and Knowledge Integration

The framework is organized around supervised, lexical, and distributional knowledge:

  1. Supervised Segmentation Layer: This layer learns word boundaries using character-level BIES tagging with CRF decoding. Baselines include maximum matching (Dictionary matching) and the linear CRF, which uses features like the current character, a two-character context window on each side, and adjacent character bigrams.

  2. Lexical Knowledge Layer: This layer encodes overlapping dictionary candidates using an aggregated lexicon-lattice representation. It derives 11 binary features (B2–B5+, I3–I5+, E2–E5+) for each character position, and estimates entry-specific reliability using a formula: r(w) = hit(w) + κpg(w) occ(w) + κ, where κ = 5 controls the smoothing strength.

  3. Distributional Knowledge Layer: This layer exploits unlabeled text in three ways: explicit distributional features (log bigram frequency, character association strength, right/left-neighbor entropy), static character embeddings (Char2Vec), and contextual pretraining via TangutEncoder.

Contextual Pretraining with TangutEncoder

TangutEncoder (TEnc) is a compact BERT-style encoder tailored to the Tangut writing system. It is pretrained with character-level masked language modeling that combines single-character and contiguous-span masking. Approximately 15% of character positions are selected using a mixture of single-character masking and contiguous spans of two to four characters, with 80% of selected positions are replaced with the mask token, 10% with random characters, and 10% remain unchanged. This model learns contextual character representations without assuming word boundaries in advance.

Knowledge Fusion and Experimental Protocol

External knowledge is incorporated at different stages: in the linear CRF, lexicon-lattice and corpus-statistical features are added directly as real-valued feature functions. In TangutEncoder, external features are incorporated after contextual encoding. The experimental protocol involves line-level five-fold crossvalidation stratified by genre to evaluate within-source generalization to held-out, thematically diverse passages rather than transfer to unseen documents. Evaluation is span-based: a predicted word is correct only when its boundaries exactly match the expert annotation. Metrics include Precision, Recall, and F1.

Comparison of Knowledge Fusion Strategies

The study compares various feature combinations across different models. The complete system, TangutEncoder + Dictall + Distall, achieved the highest overall performance: 0.917±0.003 F1 and 0.949±0.022 OOV-R in the secular genre, while the CRF remained stronger on religious texts. The results clarify that contextual pretraining substantially improves F1, OOV recall, and religious-text performance by training the entire contextual encoder rather than only providing static character embeddings. Furthermore, the complete systems remain complementary, showing that explicit lexical knowledge remains valuable after contextual pretraining.

Conclusion and Limitations

The study presents the first systematic investigation of automatic Tangut word segmentation, combining expert annotation, traditional dictionaries, and unlabeled text. The full TangutEncoder obtains the highest mean F1 (0.911) and OOV recall under within-source line-level evaluation. Limitations include a small corpus size, genre imbalance, and limited scope for transfer to unseen documents. Future work plans involve expanding expert annotation, refining the POS inventory, and exploring downstream evaluations in retrieval and translation.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be made to AI systems, categorized by the knowledge source they integrate:


) Improvements for Tangut Word Segmentation (WS) Systems:

  1. [Improvement] Implement a Reliability-Calibrated Lexicon-Lattice Representation as a core feature set for sequence labeling models (CRF/BiLSTM).

  2. [Improved System Capability] The AI system can perform more accurate word boundary detection, especially for low-frequency or out-of-vocabulary (OOV) words, by leveraging the confidence scores derived from the lexicon's reliability estimation rather than treating all dictionary matches equally.

  3. [Improvement] Integrate Explicit Distributional Features (Bigram Frequency, Character Association Strength) into the input layer of sequence models.

  4. [Improved System Capability] The system gains better contextual awareness beyond immediate neighbors, leading to more robust segmentation in rare contexts and improved performance on genre-specific texts (e.g., religious vs. secular).

  5. [Improvement] Utilize a Contextual Pretrained Character Encoder (TangutEncoder) instead of static character embeddings for Transformer-based models.

  6. [Improved System Capability] The AI system achieves superior learning of long-range dependencies and contextual meaning, leading to higher overall F1 scores and better generalization across diverse textual styles, as demonstrated by the TEnc model's 0.911 mean F1.

  7. [Improvement] Employ a Knowledge Fusion strategy that combines contextual pretraining (MLM) with the full lexicon representation (Lattice + Reliability).

  8. [Improved System Capability] The resulting system achieves state-of-the-art performance, specifically showing the highest overall F1 and OOV recall, effectively bridging the gap between unsupervised context learning and explicit lexical knowledge.

  9. [Improvement] Develop a Lexicon-Aware Continued Pretraining objective (MLM + WordRank loss) for character encoders.

  10. [Improved System Capability] The system can perform fine-grained span prediction during pretraining, which improves the encoder's ability to handle complex morphological variations and potentially leads to better downstream performance when transferred to segmentation tasks.

) General NLP/Machine Learning System Improvements (Derived from Methodology):

  1. [Improvement] Implement a Genre-Aware Evaluation Protocol using stratified cross-validation based on document genre (e.g., Buddhist scripture vs. secular text).

  2. [Improved System Capability] The system can be explicitly tuned or specialized to perform optimally on specific linguistic domains, as the performance metrics clearly show genre-level variations, allowing for targeted deployment in specialized archives or translation pipelines.

  3. [Improvement] Create a Feature Ablation Framework that systematically tests the contribution of each knowledge source (Lexicon, Distributional, Contextual) to the final model performance.

  4. [Improved System Capability] Researchers can efficiently determine whether adding more data (e.g., more unlabeled text vs. more expert annotations) yields higher returns in specific contexts, optimizing resource allocation for low-resource languages like Tangut or other historical scripts.

  5. [Improvement] Integrate OOV-Aware Evaluation Metrics (OOV-R and IV-R) into standard evaluation pipelines instead of relying solely on aggregate F1 scores.

  6. [Improved System Capability] The system can be rigorously assessed on its ability to handle novel vocabulary encountered in real-world historical or low-resource texts, providing a more honest assessment of its practical utility beyond the training set.

Abstract

Tangut is an extinct language whose script does not explicitly mark word boundaries. We present the first systematic study of Tangut word segmentation using 2,750 expert-annotated segments (31,893 tokens), traditional lexicons, and unlabeled text. Our framework combines a reliability-calibrated lexicon-lattice representation, explicit distributional statistics, and a lightweight character encoder pretrained with MLM. In within-source five-fold cross-validation, the model integrating TangutEncoder, CRF, and external features obtains the numerically highest main-system mean F1 of 0.911 and substantially improves recall beyond the labeled training vocabulary. We further evaluate document-level transfer on 479 segments (4081 tokens) from five works absent from the annotated training corpus. You can access our project at https://github.com/jiangli-va/TangutSeg.

Related papers