Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs
Glassbox AI
cs.CV, cs.AI
Submitted: 2026-08-13
Updated: 2026-10-07
Comments: Preliminary technical report. 19 pages, 8 figures, 4 algorithms
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 50/100
The gist: This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals
Terminology
Summary
This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments. Three specialized discriminators—global, hand, and head—each guide a corresponding expert branch in the generator toward a distinct visual region, enabling implicit feature specialization without explicit diversity losses. To stabilize this multi-discriminator system, whose early-phase training otherwise exhibits chaotic dynamics, we introduce a United Loss consensus mechanism that regularizes each discriminator toward the ensemble average at a 10% weight. Each branch further adopts a dual-pathway convolutional-transformer design with learnable AdaptiveFeatureFusion, balancing the stability of convolutions against the detail of windowed self-attention. The generator is trained using an alternating three-mode schedule (discriminator, holistic generation, branch-specialized generation). On a custom 156GB dataset with a filtered test set that removes easy and repetitive samples, our 0.2B-parameter variant achieves 29.8 PSNR and the 1.3B-parameter variant achieves 30.7 PSNR, with inference VRAM footprints of 1.5 GB and 8 GB respectively, enabling deployment on consumer-grade hardware. Full ablation studies remain ongoing due to the 2–3 month training cycle on a single GPU. The system was showcased at the 2025 Hong Kong Frontier Technology Summit.
The Multi-Discriminator GAN (MD-GAN) comprises three discriminators: a Global Discriminator for overall video coherence, a Hand Discriminator for gesture details, and a Head Discriminator for facial expressions, ensuring precise supervision of sign language video synthesis. The Global Discriminator processes 448×448 images across five resolution levels (448, 224, 112, 56, 28), using Haar wavelet transforms to decompose inputs into frequency subbands (6 to 24 channels). It employs MiniBatch Standard Deviation (MBSD) to prevent mode collapse and outputs a 30×30 feature map via a PatchGAN structure to enhance detail supervision. Local discriminators operate on 112×112 patches across three scales (112, 56, 28). The Hand Discriminator extracts hand regions using keypoint-based localization, scaling to 112×112 with center-crop alignment, focusing on finger positions and hand shapes. The Head Discriminator processes facial regions using nose keypoints with a 96-pixel radius, capturing expression details. Haar wavelet transforms decompose RGB images into frequency subbands, expanding the input channels from 6 to 24 (6×4 subbands) to simultaneously capture structural (low-frequency) and textural (high-frequency) features.
The generator uses a Multi parallel U-Net architecture built upon the Mixture-of-Experts (MoE) paradigm, activating all branches simultaneously for every input, ensuring uniform feature extraction coverage and balanced gradient flow. Each branch is driven by its corresponding specialized discriminator (hand, face, or global), which provides implicit pressure for the branch to specialize on specific visual regions without requiring explicit diversity losses. All three branches share the same 448 × 448 input. Each branch uses a two-stage dual-pathway architecture: Stage 1 fuses a ConvTransformerBlock stream (channel-wise MLP with LayerNorm) and a linear-projection/Attn Conv2d stream via an AdaptiveFeatureFusion (AFF) module. Stage 2 splits the fused features into a purely convolutional Attn Conv2d block and a Swin Transformer block with Window-based Multi-head Self-Attention (W-MSA), fused by a second AFF module. The AFF fusion mechanism computes a soft weighted combination of its two input streams: Fused = α · Stream1 + (1 − α) · Stream2, where α is produced by applying softmax or sigmoid to a learnable parameter. The convolutional stream produces globally stable but visually smooth features, while the Swin Transformer stream captures finer detail but exhibits instability at high-frequency boundaries—such as clothing edges, finger contours, and face-background transitions. The learnable AFF fusion resolves this tension: the convolutional stream acts as a stabilizing anchor that suppresses boundary jitter, while the Transformer stream injects the fine detail that pure convolution lacks.
The decoder employs a symmetric architecture with Upsample Vit blocks that mirror the encoder’s dual-pathway design. Each Upsample Vit block first upsamples the incoming feature map via bilinear interpolation, then concatenates the corresponding skip connection from the encoder—a concatenation of the AdaIN-fused head, hand, and global features. The combined features are split into three MoE branches (head, hand, global), each processed independently through a conv linear projection followed by parallel CNN and Swin Transformer pathways fused via AdaptiveFeatureFusion. At select layers (up 3, up 5, up 7), keypoints info f from the MappingNetwork is injected through a SpatialTransformer and CrossSwinTransformerBlock, enabling keypoint-guided cross-attention. The final block, Upsample Transformer, is a simplified single-channel variant that omits MoE branches, skip injection, and cross-attention, serving purely as a refinement stage. The MappingNetwork encodes 133 skeleton joints with 3D coordinates via a Transformer encoder (3 layers, 8 heads) with positional embeddings, producing keypoints info f shared across all decoder layers. A Local-Global Merged Attention mechanism integrates three spatial scales with weights: Merged Attention = 0.85 × Local + 0.1 × Sub-Local + 0.05 × Global. The model takes as input two three-channel images: a skeleton image and a style image, processed through separate encoders and combined using Adaptive Instance Normalization (AdaIN) at each encoder layer, facilitating the integration of structural (skeleton) and stylistic (appearance) information across multiple scales.
The United Loss function unifies the training objectives of all discriminators through a soft consensus mechanism. Each discriminator computes its own adversarial loss independently, but is additionally regularized by the ensemble average of all three discriminators at a weighting of 10% (i.e., λunited = 0.1). For each discriminator Di: Ltotal Di = Ladv Di + λunited · Lunited, where Lunited is computed from the equally-weighted average output of all three discriminators. This formulation ensures that no single discriminator can deviate too far from the ensemble consensus during the unstable early phase, while still preserving individual specialization. The 10% weight was chosen empirically to balance stability (higher weight) against discriminative diversity (lower weight). United Loss breaks the runaway feedback loop where the generator capitulates to a single discriminator: a discriminator whose individual loss has collapsed contributes a disproportionately large share of its own total loss through the consensus term, pulling the deviating discriminator back toward consensus without suppressing the legitimate specialization signals of the others.
The training loop cycles through three fixed modes in a deterministic rotation: mode(s) = ⌊s/10⌋ mod 3, yielding a 1:1:1 rotation in which each mode persists for 10 consecutive steps. Mode 0 (Discriminator mode): All three discriminators are updated in parallel while the generator is frozen. Mode 1 (Generator overall mode): The aggregated generator loss LG is backpropagated through all generator parameters simultaneously, providing holistic guidance from the combined global and local feedback. Mode 2 (Generator partial mode): Each branch (head, hand, global) is updated independently using its own discriminator-specific loss and its own optimizer, enabling region-targeted optimization without cross-branch gradient conflict. The generator loss is defined as Lgen united = λg Lglobal + λh Lhand + λf Lhead, where λg = 0.33, λh = 0.33, and λf = 0.33.
The dataset is a custom 156 GB dataset of sign language videos, each depicting a single vocabulary word with a consistent three-part structure: resting pose, sign language gesture, and return to resting pose. The resting-pose frames (easy samples) account for approximately 50% of each video’s duration on average. A filtered test set excludes all easy samples, ensuring that every reported metric targets the challenging transition and gesture frames where generation is prone to collapse or artifacts. When evaluated on the full unfiltered test set including easy samples, PSNR exceeds 36, but the filtered metric is considered far more informative. All models are trained on a single consumer GPU (NVIDIA RTX 4090, 24 GB or 48 GB variant) using the AdaBelief optimizer, with a cosine learning rate schedule with linear warmup over the first 50,000 steps. Full convergence requires approximately 5–6 million steps (2–3 months of training), with PSNR exhibiting a characteristic non-linear progression: extended plateaus punctuated by phase-transition breakthroughs, followed by terminal oscillation.
Results across three model scales (all sharing the same architecture but varying in channel width): Small (0.2B parameters) achieves 29.8 PSNR with 1.5 GB VRAM; Medium (0.66B parameters) achieves 30.4* PSNR with ∼5 GB VRAM (from an earlier dataset version v3, lower contrast; re-training on v4 is in progress and expected to yield 30.1–30.2); Large (1.3B parameters) achieves 30.7 PSNR with 8 GB VRAM. The monotonic PSNR improvement from 29.8 to 30.7 as parameters increase 6.5× (0.2B to 1.3B) suggests that the architecture effectively utilizes additional capacity rather than overfitting or suffering from optimization degradation.
Observational evidence consistent with the hypothesized consensus mechanism includes: (1) Non-linear convergence with plateau-to-breakthrough phase transitions, where breakthroughs in one region (e.g., hand clarity) tend to coincide with improvements in other regions within 50,000–100,000 subsequent steps; (2) Branch specialization without collapse—the hand branch generates sharper hand regions, the face branch produces more expressive facial details, and the global branch maintains better body coherence, driven purely by discriminator-specific loss signals; (3) Stability at scale—training remains stable across all three model scales without requiring gradient penalty, spectral normalization, or other common stabilization techniques beyond United Loss.
A controlled ablation comparing training with and without United Loss on the v4 architecture is planned but has not yet been executed due to the prohibitive training time (2–3 months per run on a single consumer GPU). The experimental design includes a baseline (with United Loss, λunited = 0.1, already completed with PSNR 29.8/30.7), an ablation (without United Loss, λunited = 0, planned), and varied λunited values (planned).
The version history documents the iterative development: v1 used AdaIN as style transfer, replacing concatenation of skeleton and style images with Adaptive Instance Normalization at each encoder layer, stabilizing finger generation. v2 added keypoint-guided generalization via a MappingNetwork (3-layer Transformer encoder) injecting 133-joint 3D keypoint coordinates into the decoder, dramatically improving generalization. v3 added a local hand discriminator operating on 112×112 hand crops, with generator loss extended to Lmain + 0.25 · Lhand, improving fine details like finger shadows and texture. v4 (current) introduced the multi-expert parallel architecture with three encoder-decoder branches, which immediately destabilized training—the three-discriminator system exhibited dynamics reminiscent of the three-body problem, directly motivating the United Loss consensus mechanism.
Limitations include: Training cost of 2–3 months on a single consumer GPU precluding comprehensive ablation studies; the contribution of United Loss not isolated on the current v4 architecture; scaling bottleneck where extending to more specialized branches would incur linear cost increases in both compute and VRAM; and metric limitations—only PSNR is currently reported, with perceptual metrics (FID, LPIPS), action accuracy, and human evaluation planned but not yet completed. Future work includes extending the loss-guided expert specialization paradigm to diffusion models via a proposed Loss-Guided MoE Diffusion with Label-Routed Experts, aiming to replace full branch activation at inference with a pre-trained label router that selects only relevant expert branches before generation begins.
Improvements for AI systems
Improvements to AI Systems:
- Multi-Discriminator Consensus Regularization for Stable Multi-Expert Training
-
Implement the United Loss mechanism (λ=0.1) in any multi-discriminator or multi-objective GAN/diffusion system to prevent mode collapse and training instability.
-
The improved system can train multiple specialized experts (e.g., for different body parts, objects, or modalities) without explicit diversity losses, avoiding chaotic early-phase dynamics and runaway feedback loops.
- Loss-Guided Expert Specialization Without Explicit Diversity Penalties
-
Use per-region discriminators (global, hand, head) to implicitly force generator branches to specialize on distinct visual regions.
-
The improved system can synthesize high-fidelity videos/images with fine-grained detail (e.g., fingers, facial expressions) while maintaining global coherence, without hand-crafted diversity terms.
- Dual-Pathway Convolutional-Transformer Fusion with Learnable Adaptive Feature Fusion (AFF)
-
Adopt the AFF mechanism (soft weighted sum of convolutional and Swin Transformer streams) to balance stability and detail.
-
The improved system can generate sharp edges (fingers, clothing contours) without boundary jitter, leveraging the stability of CNNs and the detail of self-attention.
- Alternating Three-Mode Training Schedule for Multi-Branch Generators
-
Use the deterministic mode rotation (discriminator → holistic generation → branch-specialized generation) to reduce gradient conflicts.
-
The improved system can optimize multiple generator branches independently without cross-branch interference, enabling targeted region refinement while maintaining overall quality.
- Keypoint-Guided Cross-Attention Injection via MappingNetwork
-
Inject 3D skeleton keypoints (133 joints) through a Transformer encoder into decoder layers using SpatialTransformer and CrossSwinTransformerBlock.
-
The improved system can generalize to unseen poses and gestures, using structural guidance to improve synthesis accuracy and robustness.
- Haar Wavelet Decomposition for Multi-Scale Frequency Supervision
-
Expand input channels via Haar transforms (6→24) to capture both low-frequency structure and high-frequency texture.
-
The improved system can better preserve fine textures (e.g., skin, fabric) while maintaining global shape, improving perceptual quality in video generation.
- Filtered Test Set for Challenging-Sample Evaluation
-
Exclude easy/repetitive frames (e.g., resting poses) from evaluation metrics.
-
The improved system can be benchmarked on genuinely difficult frames (transitions, gestures), providing more meaningful performance indicators and driving optimization toward failure cases.
- Scalable Architecture with Monotonic Performance Gains
-
Use the multi-expert U-Net with shared input and branch-specific discriminators, showing PSNR improves from 29.8 to 30.7 as parameters scale from 0.2B to 1.3B.
-
The improved system can efficiently utilize additional capacity without overfitting, enabling deployment on consumer hardware (1.5–8 GB VRAM) while maintaining high quality.
- Phase-Transition-Aware Training Dynamics
-
Leverage the observed plateau-to-breakthrough convergence pattern to design adaptive learning-rate schedules or early-stopping criteria.
-
The improved system can accelerate convergence by anticipating breakthroughs (e.g., hand clarity improvements leading to global gains within 50k–100k steps).
- Loss-Guided MoE Diffusion with Label-Routed Experts (Future Extension)
-
Replace full branch activation with a pre-trained label router that selects only relevant experts at inference.
-
The improved system can reduce inference cost linearly with expert count, enabling real-time generation on edge devices while retaining specialization quality.
Capabilities of the Improved AI System:
-
Generates realistic sign language videos with precise hand gestures, facial expressions, and body coherence.
-
Trains stably on single consumer GPUs without gradient penalties or spectral normalization.
-
Scales to larger models with predictable quality gains and low VRAM footprint.
-
Evaluates robustly on challenging frames, avoiding inflated metrics from easy samples.
-
Extends to other multi-region generation tasks (e.g., human avatars, medical imaging, autonomous driving scenes) requiring specialized local fidelity.
Abstract
This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments. Three specialized discriminators -- global, hand, and head -- each guide a corresponding expert branch in the generator toward a distinct visual region, enabling implicit feature specialization without explicit diversity losses. To stabilize this multi-discriminator system, whose early-phase training otherwise exhibits chaotic dynamics, we introduce a United Loss consensus mechanism that regularizes each discriminator toward the ensemble average at a 10% weight. Each branch further adopts a dual-pathway convolutional-transformer design with learnable AdaptiveFeatureFusion, balancing the stability of convolutions against the detail of windowed self-attention. The generator is trained using an alternating three-mode schedule (discriminator, holistic generation, branch-specialized generation). On a custom 156GB dataset with a filtered test set that removes easy and repetitive samples, our 0.2B-parameter variant achieves 29.8 PSNR and the 1.3B-parameter variant achieves 30.7 PSNR, with inference VRAM footprints of 1.5 GB and 8 GB respectively, enabling deployment on consumer-grade hardware. Full ablation studies remain ongoing due to the 2-3 month training cycle on a single GPU. The system was showcased at the 2025 Hong Kong Frontier Technology Summit.
Sources
- Wasserstein GAN
- MCL-GAN: Generative Adversarial Networks with Multiple Specialized Discriminators
- Generative Adversarial Networks
- A Style-Based Generator Architecture for Generative Adversarial Networks
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models