Federated Multilingual Speech-LLMs: Architecture and Aggregation Strategy Benchmarking
cs.CL, cs.AI
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: Accepted Iberspeech 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: We present a comprehensive benchmark of Federated Learning (FL) for multilingual Automatic Speech Recognition (ASR), evaluating four Speech-LLM architectures on the Multilingual LibriSpeech dataset.
Terminology
Abstract
We present a comprehensive benchmark of Federated Learning (FL) for multilingual Automatic Speech Recognition (ASR), evaluating four Speech-LLM architectures on the Multilingual LibriSpeech dataset. We compare FedAvg and FedProx across frozen and unfrozen encoder configurations, demonstrating that optimized learning rates are critical for performance. Specifically, independently tuning the learning rates for the speech encoder, connector, and decoder yields the lowest error rates, with full three-component adaptation (LoRA for encoder and decoder, full training for the connector) producing the best FL results. We observe that FedProx efficacy is architecture-dependent, providing notable advantages in multilingual pre-trained architectures (e.g., EuroLLM over TinyLlama when keeping the encoder fixed); this indicates that LLM backbone capacity plays a key role in mediating resilience to heterogeneous data distributions. These findings offer concrete design guidance for deploying multilingual Speech-LLMs in privacy-sensitive, distributed environments.
Sources
- SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
- EuroLLM: Multilingual Language Models for Europe
- Parameter-Efficient Transfer Learning under Federated Learning for Automatic Speech Recognition
- Federated Learning with Non-IID Data
- LoRA: Low-Rank Adaptation of Large Language Models
- TinyLlama: An Open-Source Small Language Model
- Voxtral
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering