TransSLR: A Lightweight Transformer for Sign Language Recognition

arXiv:2608.06407 · cs.CV, cs.AI · Submitted 2026-08-03 · Read on arXiv

Carnegie Mellon University Africa

cs.CV, cs.AI

Submitted: 2026-08-03

Updated: 2026-09-27

Comments: This paper has been accepted for oral presentation at Deep Learning Indaba 2026 which will be hosted on the IJCAI hosting platform. paper link: https://chairingtool.com/conferences/dli2026/main-track/submissions/356

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 49/100

The gist: The paper "TransSLR: A Lightweight Transformer for Sign Language Recognition" addresses the challenge that "Automated Sign Language Recognition for underrepresented languages remains a largely

Terminology

Summary

The paper TransSLR: A Lightweight Transformer for Sign Language Recognition addresses the challenge that Automated Sign Language Recognition for underrepresented languages remains a largely unsolved problem, specifically focusing on Central African Sign Language (CASL). The authors observe that the common heuristic of fine-tuning high-resource models fails to close [the accuracy gap] for CASL because of the limited scale of available CASL data and the significant lexical and visual domain gap between CASL and large-scale corpora such as WLASL, which renders pre-trained representations largely uninformative.

To address this, the authors propose TransSLR, a lightweight Temporal Transformer Encoder trained from scratch on 64-frame normalized pose sequences, with average pooling and a classification head.

Methodology

The methodology relies on a pipeline that transforms raw RGB footage into standardized kinematic tensors. Key components include:

  • Data Preprocessing: The system uses a uniform sampling strategy, compressing or expanding each video into a fixed 64-frame window via linear interpolation. Spatial features are extracted using the MediaPipe Holistic framework, yielding a (64, 225) matrix representing 75 skeletal keypoints. To mitigate bias, a dual-normalization protocol is applied: spatial centering by setting the nasal landmark P nose (x, y, z) as the coordinate origin and shoulder-distance scaling. Additionally, Gap-Filling logic uses temporal linear interpolation to restore missing manual landmarks.

  • Architecture: TransSLR is an encoder-only model paired with a Global Average Pooling classification head. The pipeline consists of a linear input projection, sinusoidal positional encoding, a Transformer encoder stack [comprising N=4 stacked layers], and a classification head. The encoder uses Multi-Head Self-Attention (MHSA) followed by a position-wise FeedForward Network (FFN), with residual connections and layer normalization.

Experimental Results

Evaluated on the CASL-W60 benchmark, TransSLR establishes a new state-of-the-art accuracy of 80.39%, surpassing the prior best by +10.46%. The model also achieves a Top-5 accuracy of 91.07% and a classification error rate [of] 19.61%.

The researchers found that geometric pose-only modeling significantly outperforms multimodal RGB-pose fusion for signer-independent CASL recognition. They attribute the failure of multimodal baselines (such as the MViT V2 + Bi-LSTM, which achieved a Recall@1 of only 29.90%) to biometric noise overfitting, where Raw RGB pixels allow the model to 'cheat' during the optimization phase by memorizing subject-specific artifacts such as clothing patterns, skin tone, and localized lighting conditions. In contrast, TransSLR's geometric abstraction effectively strips away the signer’s identity, forcing the Transformer Encoder to learn scale-invariant motion primitives.

Efficiency and Limitations

TransSLR is highly efficient, possessing the smallest model size, requiring only 8.67M trainable parameters, while also exhibiting the lowest computational cost at 0.277 GFLOPs per 64-frame pose sequence. This is compared to VideoMAE, which requires 65.02M [parameters] and 101.85 [GFLOPs].

Regarding limitations, the authors identified Linguistic Dead Zones where certain classes (30, 35, 51, 57, and 59) returned a 0.00% F1-score. This is attributed to High-Frequency Motion Overlap, where signs share identical manual trajectories and differ only in non-manual markers, suggesting that incorporating facial or non-manual features may improve discrimination in these cases.

Improvements for AI systems

1. Multi-Stream Non-Manual Feature Fusion

  • The Improvement: Integrate a secondary, specialized Transformer stream dedicated to facial landmarks (using MediaPipe Face Mesh) that interacts with the existing body/hand pose stream via cross-attention mechanisms.

  • What the improved system can do: It can resolve Linguistic Dead Zones by distinguishing between signs that share identical hand trajectories but differ in facial expressions, mouthings, or head tilts, thereby eliminating the 0.00% F1-score errors in morphologically similar sign classes.

2. Self-Supervised Masked Pose Modeling (MPM) Pre-training

  • The Improvement: Implement a self-supervised pre-training phase where the Transformer is tasked with reconstructing randomly masked skeletal keypoints from large-scale, unlabeled sign language video datasets.

  • What the improved system can do: It can overcome the limited scale of available CASL data by learning universal motion primitives and spatial dependencies from unlabeled video, allowing the model to achieve high accuracy in underrepresented languages with significantly fewer labeled examples.

3. Multi-Scale Temporal Hierarchical Encoding

  • The Improvement: Replace the fixed 64-frame linear interpolation with a hierarchical temporal encoder that utilizes dilated temporal convolutions or multi-scale sampling rates.

  • What the improved system can do: It can capture both high-frequency rapid manual gestures and low-frequency subtle body/postural shifts simultaneously, preventing information loss caused by compressing variable-speed signing into a rigid 64-frame window.

4. Cross-Lingual Knowledge Distillation

  • The Improvement: Use a heavy, multimodal Teacher model (trained on high-resource RGB datasets like WLASL) to supervise the lightweight TransSLR Student via a distillation loss.

  • What the improved system can do: It allows the lightweight pose-only model to inherit complex semantic features from high-resource RGB models without inheriting the biometric noise (clothing, skin tone, lighting) that causes overfitting, resulting in a model that is both computationally tiny and highly generalized.

Sources

Related papers