A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language

arXiv:2608.10588 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Ushnish Sarkar, Suvajit Patra, Bhaswar Chattopadhyay, Pranab Singha Roy, Tapas Samanta

Variable Energy Cyclotron Centre · Homi Bhabha National Institute · Ramakrishna Mission Vivekananda Educational and Research Institute

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 50/100

Terminology

Summary

Summary

This paper introduces a new dataset and baseline models for fine-grained isolated handshape recognition in sign language, grounded in the Hamburg Notation System (HamNoSys). The work addresses the need for a real-image benchmark that combines a broad, transcription-defined handshape inventory with systematic evaluation on unseen participants.

Dataset Construction: The dataset comprises 144,000 RGB images collected from 15 participants (university students aged 23–25 years) for 160 handshape classes defined by the official, non-exhaustive HamNoSys 4 Handshapes Chart. Each distinct illustrated hand model in the chart was treated as one class, with blank cells and cells containing only symbols or cross-references excluded. The 160 classes are grouped into ten populated categories: six Selection categories (Fist, One Finger, Two Fingers (nonspread), Two Fingers (spread), Flathand (Four Fingers nonspread), Four Fingers (spread)) and four Thumb-opposition categories (One Finger others in fist position, Two Fingers (nonspread) others in fist position, Four Fingers (nonspread), One Finger others extended (spread)). The classes are also grouped by chart column: 106 Selection classes and 54 Thumb-opposition classes.

Acquisition Protocol: Recordings were made indoors against a constant dull-white background using a tripod-mounted Logitech Brio RGB camera at 30 frames/s and 640 × 480 pixels. Each participant performed every class, maintaining the handshape while slowly rotated the hand about two approximately orthogonal axes to introduce viewpoint and self-occlusion variation. Each 10-second clip yielded 300 frames, from which every fifth frame was selected, producing 60 images per clip. MediaPipe Hands localised the dominant hand in 139,199 images (96.66% of the total), which formed the common modelling subset for all baselines; the remaining 4,801 images were retained in the complete dataset.

Evaluation Protocols: Two complementary protocols were used: (1) a class-stratified subject-dependent split with a 70:15:15 train/validation/test partition at frame level, and (2) a 15-fold leave-one-subject-out (LOSO) protocol where one participant was held out for testing in each fold, with the remaining 14 participants' images divided into class-stratified training and validation partitions using an 85:15 ratio.

Baseline Models: Four models were evaluated: two appearance-based models (ResNet-18 and ViT-B/16) operating on RGB hand crops, and two landmark-based models (a graph convolutional network and XGBoost) operating on MediaPipe hand landmarks. For the landmark models, 21 hand landmarks were extracted, with coordinates translated, rotated into a hand-centred frame, and scaled; a joint-angle feature normalised by π was computed for the 15 internal finger joints.

Subject-Dependent Results: On the proposed 160-class dataset, the highest top-1, top-3, and top-5 accuracies were achieved by ViT-B/16, at 86.20%, 95.99%, and 97.79%, respectively. ResNet-18 achieved a comparable top-1 accuracy of 84.72%. The GCN achieved 72.44% top-1 accuracy, and XGBoost achieved 69.57%. The higher top-1 accuracies of the RGB-based models are consistent with fine-grained distinctions benefiting from appearance information that is not completely retained by the 21-point landmark representation. Weighted and macro scores were closely aligned, consistent with the approximately balanced class distribution.

Confusion Analysis: The broad-category confusion matrix for ViT-B/16 is strongly concentrated along the diagonal, but off-diagonal predictions occurred primarily between related categories. The five most frequently confused target-level pairs included FTR OFOE2/FTR OFOE3 (pair error 12.26%), FTI TFO1/FTR TFO1 (11.99%), SFB FFS1/SFB FFS3 (11.92%), FTR OFO2/FTR OFO3 (10.94%), and DE F5/SFE F1 (9.36%). These pairs differ only in subtle properties such as finger selection, bending, thumb position, or contact.

LOSO Results: On the proposed dataset, numerically close mean top-1 accuracies were obtained by ResNet-18 and ViT-B/16, at 45.38% and 45.22%, respectively. The GCN achieved 43.49% and XGBoost 39.66%. The GCN achieved the highest top-3 and top-5 accuracies at 69.23% and 78.58%. Relative to the subject-dependent evaluation, the ResNet-18 and ViT-B/16 top-1 accuracies decreased by 39.34 and 40.98 percentage points, respectively, with substantial fold standard deviations. These results identify unseen-participant generalisation as the principal challenge of the proposed 160-class benchmark.

External References: Matched-model evaluations were conducted on LSWH100 (100 synthetic SignWriting-derived classes) and ASL Fingerspelling Dataset A (24 static ASL letters). On LSWH100, ResNet-18 achieved the highest top-1 accuracy at 86.70%. On ASL Fingerspelling Dataset A, top-1 accuracy above 97% was obtained by all four models under subject-dependent evaluation, and mean LOSO top-1 accuracy ranged from 82.20% for ViT-B/16 to 87.40% for ResNet-18. These external results provide context rather than direct estimates of relative dataset quality due to different class inventories and acquisition conditions.

Conclusion and Limitations: The dataset and baselines are intended to support phonology-grounded research and the development of transcription, recognition, and translation tools across sign languages, including under-resourced settings. Limitations noted include: data collected from only 15 university students in a controlled indoor environment with a single RGB camera; the class inventory restricted to 160 static, single-hand forms from the non-exhaustive chart (excluding dynamic transitions, two-handed configurations, orientation, location, movement, and non-manual components); operator verification rather than validation by HamNoSys specialists; and evaluation of only four baseline model families.

Improvements for AI systems

Improvements to AI Systems:

  1. Fine-grained handshape recognition with explicit phonological grounding – Train models to predict HamNoSys feature vectors (e.g., finger selection, spread, thumb opposition) as intermediate or auxiliary outputs, enabling compositional generalization to unseen handshapes not in the 160-class inventory.

  2. Viewpoint-invariant hand pose encoding – Use the rotation-augmented training data (slow hand rotation about two orthogonal axes) to train a self-supervised view-invariant encoder, improving robustness to self-occlusion and arbitrary camera angles in real-world sign language capture.

  3. Cross-subject generalization via domain adaptation – Leverage the 15-fold LOSO protocol to train a meta-learning or adversarial domain-adaptation module that aligns hand appearance and landmark distributions across participants, reducing the observed 40-point accuracy drop for unseen users.

  4. Hybrid RGB-landmark fusion – Combine appearance (ViT/ResNet) and landmark (GCN) streams with a cross-attention mechanism, using the landmark graph to guide spatial attention on finger joints and contact points, targeting the identified confusion pairs (e.g., FTR OFOE2/FTR OFOE3) that differ only in subtle bending or contact.

  5. Phonological confusion-aware loss – Modify the training loss to penalize errors between HamNoSys-related classes more heavily (e.g., using a hierarchical loss based on the chart’s Selection and Thumb-opposition categories), improving discrimination among the five most confused pairs.

  6. Temporal augmentation for static-shape extraction – Use the 60-frame-per-clip sampling to train a temporal consistency module that predicts a single handshape from multiple frames, filtering out motion blur and transient occlusion, improving accuracy on low-quality or partially visible hands.

  7. Few-shot and zero-shot handshape recognition – Use the HamNoSys transcription as a semantic embedding space to enable zero-shot classification of new handshapes (e.g., from the non-exhaustive chart) by mapping unseen feature combinations to known visual prototypes.

  8. Uncertainty-aware landmark refinement – Train the GCN to output per-landmark confidence scores, then use these to weight the fusion with RGB features, reducing the impact of the 3.34% of images where MediaPipe failed to localize the hand.

What the improved AI system can do:

  • Recognize handshapes from unseen signers with significantly higher accuracy (targeting >60% top-1 in LOSO, up from 45%).

  • Generalize to novel handshapes defined by HamNoSys features without retraining.

  • Operate robustly in varied lighting, backgrounds, and camera angles due to viewpoint-invariant encoding.

  • Provide interpretable predictions by outputting phonological feature labels (e.g., “fist, thumb opposed, index extended”), aiding transcription and translation tools for under-resourced sign languages.

  • Distinguish subtle handshape pairs (e.g., bent vs. straight fingers, thumb contact vs. non-contact) that current models confuse, improving fine-grained sign language understanding.

Abstract

Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited. This work introduces a benchmark grounded in the language-independent Hamburg Notation System (HamNoSys). Methods: A balanced dataset of 144,000 RGB images was collected from 15 participants for 160 handshape classes defined by the official HamNoSys 4 Handshapes Chart. ResNet-18 and ViT-B/16 were evaluated as appearance-based models, while a graph convolutional network and XGBoost were evaluated from hand landmarks. Both a class-stratified subject-dependent split and a 15-fold leave-one-subject-out (LOSO) protocol were used. The same model families were additionally assessed on LSWH100 and ASL Fingerspelling Dataset A for external context. Results: The subject-dependent benchmarks established reproducible reference performance across all four model families, whereas LOSO evaluation exposed a substantial reduction when recognition was required to generalise to unseen participants. On ASL Fingerspelling Dataset A, mean LOSO top-1 accuracy ranged from 82.20% to 87.40%. Conclusion: The documented acquisition, curation, and complementary evaluation protocols pro-vide a reproducible resource for fine-grained isolated-handshape research and for developing more accessible sign-language technologies.

Sources

Related papers