Donors and Recipients: On Asymmetric Transfer Across Tasks and Languages with Parameter-Efficient Fine-Tuning
cs.CL, cs.AI
Submitted: 2025-11-17
Updated: 2026-09-11
Code: https://github.com/huggingface/peft
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) perform strongly across tasks and languages, yet how improvements in one task or language affect other tasks and languages remains poorly understood.
Terminology
Abstract
Large language models (LLMs) perform strongly across tasks and languages, yet how improvements in one task or language affect other tasks and languages remains poorly understood. We conduct a controlled LoRA fine-tuning study across multiple open-weight LLM families and scales, using a standardised grid of 11 languages and four benchmarks. We fine-tune each model on a single task-language source, then evaluate it on all other task-language target pairs to measure transfer. We decompose transfer into three regimes: (i) Matched-Task (Cross-Language), (ii) Cross-Task (Matched-Language), and (iii) Cross-Task (Cross-Language). Single-source fine-tuning yields a net positive uplift across regimes, but the gains are strongly asymmetric. Matched-Task (Cross-Language) transfer emerges as the most effective and structurally regular regime, with transfer magnitude driven principally by the identity of the target language rather than model architecture. We identify a stable coarse-grained hierarchy in which some task and language targets consistently absorb gains from diverse sources, while others remain relatively isolated. These results imply that effective fine-tuning requires accounting for donor-recipient roles to maximise downstream gains while limiting collateral degradation in other capabilities.
Sources
- Zero-shot cross-lingual transfer in instruction tuning of large language models
- Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- The Llama 3 Herd of Models
- An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks
- Language Models' Factuality Depends on the Language of Inquiry
- Qwen2.5 Technical Report
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
- Gemma 3 Technical Report
- Finetuned Language Models Are Zero-Shot Learners
- Exploring Accuracy-Fairness Trade-off in Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering