Gloss-Free Representation Learning for Cross-Dataset Sign Spotting

arXiv:2608.11332 · cs.CL, cs.CV · Submitted 2026-08-11 · Read on arXiv

Oğuz Akif Tüfekcioğlu, Ezgi Ekin, Mustafa Kaan Çevik, Hacer Yalim Keles

Hacettepe University

cs.CL, cs.CV

Submitted: 2026-08-11

Updated: 2026-08-13

Comments: Accepted at the 4th LIMIT Workshop (Representation Learning with Very Limited Resources), ECCV 2026. The abstract was shortened to comply with arXiv's 1,920-character limit

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 75/100

The gist: Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels, including glosses, temporal boundaries, and sign order.

Terminology

Summary

Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels, including glosses, temporal boundaries, and sign order. Broadcast news provides a practical alternative by pairing continuous signing with spoken-language transcripts, but this supervision is weak because text and signing are only loosely aligned. Morphologically rich languages such as Turkish add a further difficulty, since the same lexical meaning can appear in many inflected forms, while some derived forms should remain separate. This paper studies whether weak transcript-based supervision can pretrain a reusable sign encoder in this morphologically rich setting, where inadequate text normalization can fragment pseudo-gloss targets and weaken representation learning. Unlike prior pseudo-gloss pipelines designed mainly to improve translation, the authors test whether the pretrained encoder transfers as a reusable representation for cross-dataset sign spotting.

The authors pretrain on TSL-News, a Turkish broadcast corpus collected for this work, using pseudo-gloss labels derived from transcripts rather than manually annotated glosses. They compare two pseudo-gloss construction strategies: rule-based morphological lemmatization and constrained LLM-assisted normalization over a fixed vocabulary. The learned representations are evaluated through cross-dataset sign spotting on a new TSL Spotting Benchmark (TSL-SB) built from the TSL Dictionary corpus (TSLD). The LLM-assisted encoder raises top-5 temporal localization mean IoU from 0.235 with raw spatial features to 0.465, with 56.2% of examples reaching an IoU of at least 0.50. A frequency analysis further suggests that localization quality is not mainly driven by memorization of frequent pseudo-gloss labels. In a downstream translation check, the same pretraining improves BLEU-4 from 9.60 to 11.04 and ROUGE from 23.48 to 27.43. These results show that loosely aligned broadcast data can provide effective weak supervision for learning sign representations that capture both lexical content and temporal structure.

The paper makes four contributions. First, it repurposes a Sign2GPT-style pseudo-gloss pretraining pipeline – originally proposed to supply weak supervision for gloss-free translation – to ask whether the resulting representation is itself strong enough to support cross-dataset sign spotting, a task central to both sign language recognition and translation. This reframes an existing weak-supervision pretraining paradigm as a representation-learning question, tested outside both its original translation objective and its training corpus. Second, it designs two Turkish-specific pseudo-gloss construction strategies, moving from rule-based morphological lemmatization to a constrained LLM-assisted lexical normalization, to address the lexical variation introduced by Turkish's agglutinative morphology. Third, it documents TSL-News and TSL-SB as broadcast pretraining and cross-dataset spotting resources for this setting. Fourth, it introduces a multi-angle evaluation protocol – combining isolated-template NCC matching, temporal IoU, direct pseudo-gloss score localization, and all-vocabulary local NCC – that isolates representation quality from confounds such as vocabulary coverage or classifier bias, complemented by a downstream translation check.

The method uses a stage-1 model that maps a sign-language video into temporal hidden states via a trainable video encoder. The encoder uses DINOv2 for frame features with lightweight adaptation, followed by a MetaFormer temporal encoder for token mixing with local attention and downsampling. The output hidden states are passed to a prototype head whose class prototypes are initialized from Turkish fastText subword word embeddings. The model is trained with binary cross entropy over the sentence-level pseudo-gloss set. Because no temporal target is supplied, any temporal structure in the hidden states must emerge from video dynamics, the encoder inductive bias, and the pressure to explain the weak pseudo-gloss set.

For Turkish pseudo-gloss construction, the morphology-lemma strategy uses rule-based Turkish morphological analysis, returning root morphemes as lemma candidates, yielding a pseudo-gloss vocabulary of 4802 classes. The LLM lexical-rule strategy uses an LLM as a constrained lexical normalizer, selecting content-bearing Turkish units while suppressing function words, removing inflectional morphology, and preserving derivational suffixes when they change lexical meaning. The resulting pseudo-gloss vocabulary contains 6,539 classes, 1737 more than the morphology-lemma vocabulary.

TSL-News contains TV news broadcasts from 2021–2023 with Turkish transcripts but no manual gloss labels, sign-order labels, or temporal sign boundaries. It has 13,378 sentence segments (11,507 train / 802 validation / 1,069 test), 21+ hours of video, 3 signers, and sentence-level supervision. TSL-SB contains 596 annotated sign words and 1842 sentence-level temporal annotations, of which 1817 fall inside the LLM lexical-rule vocabulary and 1159 inside the morphology-lemma vocabulary, with a shared intersection of 1137 used for controlled comparison. The benchmark contains 4363 videos total.

The evaluation uses normalized cross-correlation (NCC) between hidden-state sequences. Given an isolated TSLD sign video and a continuous TSLD sentence video, the isolated-sign representation is slid over the continuous representation, and the highest-scoring temporal windows are compared with human temporal annotations. All main runs trim 0.5 seconds from both ends of the continuous search region and 0.2 seconds from both ends of isolated signs.

In the main NCC results, the spatial-feature baseline achieves top-1/top-3/top-5 mean IoU of 0.124/0.192/0.235 with 23.9% Top-5@0.50. The morphology-lemma encoder achieves 0.173/0.293/0.363 with 42.8% Top-5@0.50 on 1159 examples. The LLM lexical-rule encoder achieves 0.271/0.403/0.465 with 56.2% Top-5@0.50 on 1817 examples. On the shared 1137-example subset, the morphology-lemma encoder reaches 0.175/0.297/0.368, while the LLM lexical-rule encoder reaches 0.271/0.412/0.473, corresponding to gains of 0.096/0.115/0.105.

The LLM lexical-rule strategy increases the matched TSLD sign vocabulary from 1,060 to 1,398 signs and reduces the number of benchmark examples skipped due to vocabulary mismatch from 683 to 25. The top-5 IoU survival curves show that the spatial baseline completely misses the sign (top-5 IoU = 0) in about half of examples, while the LLM lexical-rule encoder cuts the complete-miss rate to about one quarter. The median top-5 IoU rises from 0.000 for the spatial baseline and 0.358 for the morphology-lemma encoder to 0.553 for the LLM lexical-rule encoder.

A frequency analysis relating localization quality to the number of TSL-News train sentences in which each target pseudo-gloss appears shows a weak association (Spearman rho = -0.069), suggesting that the LLM lexical-rule encoder is not simply memorizing frequent pseudo-gloss labels. The benchmark contains only signs that appear at least once in the pretraining text, so truly unseen signs are not tested, but within this matched range, localization quality does not depend on how often a sign was seen during pretraining.

In the pseudo-gloss pretraining validation, the raw surface-token baseline achieves a validation F1 of 0.1424, the LLM lexical-rule achieves 0.4732 (+0.3308 gain), and the rule-based morphology lemmas achieve 0.4791 (+0.3367 gain). In the translation check, without pseudo-gloss pretraining the model achieves BLEU-1/2/3/4 of 23.41/16.44/12.27/9.60 and ROUGE of 23.48. With LLM lexical-rule pretraining, BLEU-1/2/3/4 become 27.02/19.00/14.03/11.04 and ROUGE becomes 27.43. With morphology-lemma pretraining, BLEU-1/2/3/4 become 29.25/21.06/15.93/12.65 and ROUGE becomes 28.87.

Auxiliary localization diagnostics on the LLM lexical-rule benchmark show that the pseudo-gloss classifier achieves top-1/top-3/top-5 IoU of 0.241/0.279/0.280 with 28.9% Top-5@0.50 on 1817 examples. Target-known NCC achieves 0.271/0.403/0.465 with 56.2% Top-5@0.50 on 1817 examples. All-vocabulary local NCC achieves 0.377/0.482/0.516 with 56.3% Top-5@0.50 on 2071 segment occurrences. The all-vocabulary local NCC retrieval metrics show MRR of 0.212, Recall@1 of 13.4%, Recall@5 of 29.1%, Recall@10 of 36.0%, median rank of 33, and average candidates of 2491.

The paper concludes that even noisy sentence-level pseudo-glosses can shape a visual encoder into a lexical-temporal representation that supports dictionary-query sign spotting and provides a useful starting point for downstream translation checks. The evaluation remains limited to representation quality through sign spotting: TSL-SB relies on Turkish word-form overlap rather than expert sign glosses, isolated dictionary productions can differ from sentence-context productions, TSL-News contains only three signers, and the all-vocabulary diagnostic still searches within the annotated interval plus a small buffer. Future work will extend this diagnostic toward open-ended proposals, sentence-level reranking, broader TSL-SB validation, and full translation studies. The TSL-News corpus and the TSL-SB benchmark annotations will be released upon publication.

Improvements for AI systems

Improvements to AI systems:

  1. Weakly-Supervised Sign Encoder Pretraining
  • Train a reusable sign-language video encoder using only loosely aligned broadcast transcripts (no manual glosses or temporal boundaries), enabling scalable pretraining for low-resource sign languages.

  • The encoder learns lexical and temporal structure from sentence-level pseudo-glosses, producing hidden-state representations that transfer across datasets.

  1. Morphology-Aware Pseudo-Gloss Generation
  • Use a constrained LLM-assisted lexical normalizer (instead of rule-based lemmatization) to handle agglutinative languages like Turkish, reducing vocabulary fragmentation and improving pseudo-gloss consistency.

  • This increases matched sign vocabulary (e.g., from 1,060 to 1,398 signs) and reduces skipped benchmark examples (from 683 to 25).

  1. Cross-Dataset Sign Spotting via NCC Matching
  • Implement normalized cross-correlation (NCC) between isolated sign templates and continuous sentence representations for dictionary-query spotting, achieving top-5 mean IoU of 0.465 (vs. 0.235 with raw features) and 56.2% of examples reaching IoU ≥ 0.50.

  • Enable open-vocabulary retrieval by sliding templates over full sequences without requiring temporal annotations.

  1. Robust Evaluation Protocol for Representation Quality
  • Use multi-angle diagnostics (target-known NCC, all-vocabulary local NCC, pseudo-gloss score localization) to isolate representation quality from classifier bias or vocabulary coverage.

  • All-vocabulary local NCC achieves MRR of 0.212 and Recall@5 of 29.1%, showing the encoder supports realistic search over thousands of candidate signs.

  1. Downstream Translation Improvement
  • Use the pretrained encoder as initialization for sign-to-text translation, improving BLEU-4 from 9.60 to 11.04 (LLM lexical-rule) and to 12.65 (morphology-lemma), with ROUGE rising from 23.48 to 27.43–28.87.
  1. Frequency-Robust Representation Learning
  • The encoder does not simply memorize frequent pseudo-gloss labels (Spearman rho = -0.069 between localization quality and label frequency), indicating it learns generalizable sign patterns rather than dataset-specific shortcuts.

What the improved AI system can do:

  • Sign spotting in continuous video from a dictionary query, even when the query sign appears in varied sentence contexts, with high temporal localization accuracy (median top-5 IoU = 0.553).

  • Pretrain on any broadcast corpus with transcripts alone, eliminating the need for expensive manual glossing, making it feasible for resource-constrained sign languages.

  • Handle morphologically rich languages by normalizing inflected forms into consistent lexical units, improving vocabulary coverage and reducing false negatives.

  • Transfer across datasets (e.g., from TSL-News to TSL-SB) without fine-tuning, demonstrating reusable representations.

  • Support both recognition and translation tasks: the same encoder powers sign spotting and improves downstream translation quality, serving as a general-purpose visual backbone for sign-language AI.

  • Scale to open-vocabulary search over thousands of signs (all-vocabulary NCC with 2,491 average candidates), enabling practical dictionary lookup and retrieval in real-world applications.

Abstract

Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels such as glosses, temporal boundaries, and sign order. Broadcast news offers a practical alternative by pairing continuous signing with spoken-language transcripts, but this supervision is weak since text and signing are loosely aligned. Morphologically rich languages such as Turkish add further difficulty, as the same lexical meaning can appear in many inflected forms while some derived forms should remain distinct. We study whether weak transcript-based supervision can pretrain a reusable sign encoder in this setting, where poor text normalization can fragment pseudo-gloss targets and weaken representation learning. Unlike prior pseudo-gloss pipelines designed mainly to improve translation, we test whether the pretrained encoder transfers as a reusable representation for cross-dataset sign spotting. We pretrain on TSL-News, a new Turkish broadcast corpus, using pseudo-gloss labels derived from transcripts rather than manual annotation, comparing rule-based morphological lemmatization with constrained LLM-assisted normalization over a fixed vocabulary. We evaluate the learned representations via cross-dataset sign spotting on a new TSL Spotting Benchmark built from the TSL Dictionary corpus. The LLM-assisted encoder raises top-5 temporal localization mean IoU from 0.235 to 0.465, with 56.2% of examples reaching an IoU of at least 0.50; a frequency analysis suggests this gain is not mainly driven by memorizing frequent pseudo-gloss labels. In a downstream translation check, the same pretraining improves BLEU-4 from 9.60 to 11.04 and ROUGE from 23.48 to 27.43. These results show that loosely aligned broadcast data can provide effective weak supervision for learning sign representations that capture both lexical content and temporal structure.

Sources

Related papers