Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning

arXiv:2609.12584 · cs.LG, cs.AI · Submitted 2026-09-11 · Read on arXiv

cs.LG, cs.AI

Submitted: 2026-09-11

Updated: 2026-09-11

Code: https://github.com/kaist-dmlab/CluSTER

License: http://creativecommons.org/licenses/by/4.0/

The gist: Instruction-tuning datasets for large language models (LLMs) are often large, redundant, and imbalanced, limiting efficient adaptation.

Terminology

Abstract

Instruction-tuning datasets for large language models (LLMs) are often large, redundant, and imbalanced, limiting efficient adaptation. Naive large-batch fine-tuning repeatedly includes overrepresented sample groups while weakly covering underrepresented but informative ones, especially under data parallelism (DP) across multiple GPUs. We propose CluSTER, a Cluster-aware balanced Sampling framework for Training Efficient data Reduction in DP instruction tuning. CluSTER curates a representative reduced dataset through gradient-space clustering and DP-aware balanced allocation, ensuring dual-level coverage across clusters and workers, while preserving the original data distribution by weighted update. As a result, CluSTER reduces redundant computation and improves training stability without compromising model quality. Across multiple instruction-tuning datasets, CluSTER reduces training time by up to 69.6% with almost no accuracy loss compared to prior sampling and data reduction methods. Code is available at https://github.com/kaist-dmlab/CluSTER.

Related papers