A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models
Wajdi Ben Saad, Safa Madiouni
Carthago Labs · Université Paris Dauphine-PSL
cs.CL, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: Accepted for publication at the 16th International Conference on Advanced Computer Information Technologies (ACIT 2026), https://acit.tech/
Code: https://github.com/WajdiBenSaad/multilingual-routing-classifier
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 75/100
The gist: This paper presents a cost-efficient routing pipeline for multilingual short-text classification using small language models, evaluated on two benchmarks: a 15-language subset of SIB-200 for
Terminology
Summary
This paper presents a cost-efficient routing pipeline for multilingual short-text classification using small language models, evaluated on two benchmarks: a 15-language subset of SIB-200 for seven-way topic classification and a 15-locale subset of MASSIVE for intent classification over an official 60-intent inventory.
The pipeline is fully self-hosted, uses pretrained compact sentence encoders, and requires no task-specific fine-tuning. It evaluates a fixed-list routing strategy that keeps stronger languages on a direct multilingual path and selectively sends weaker languages through translation into English before zero-shot classification. The multilingual path uses sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2, while the translate-then-classify path uses sentence-transformers/paraphrase-MiniLM-L6-v2. Translation uses Helsinki-NLP OPUS-MT models where a dedicated source-to-English checkpoint exists, with a fallback to facebook/nllb-200-distilled-600M otherwise.
The router is a static lookup from language to inference path, with languages assigned to three analysis tiers (high, mid, low). Four routing configurations are evaluated: R0 (no translation), R1 (translate only the low-tier), R2 (translate mid-tier and low-tier), and R3 (translate all tiers). The multilingual-only baseline is equivalent to R0, and the translation-only baseline is equivalent to R3.
The main finding is that the best routing boundary is not universal. On SIB-200, the best overall configuration is R1, which translates only the low-resource tier: high-tier and mid-tier Macro-F1 remain unchanged, while low-tier Macro-F1 rises from 0.4632 to 0.6828, a gain of +0.2196. R1 achieves overall Macro-F1 of 0.7403 and accuracy of 0.7415. Moving the boundary to R2 lowers Macro-F1 to 0.7195 while increasing mean latency from 0.27534 s to 0.40952 s. Full translation (R3) remains below R1 on overall quality.
On the MASSIVE subset, the same low-tier intervention under R1 raises low-tier Macro-F1 from 0.2143 to 0.4417, a gain of +0.2275, while high-tier and mid-tier remain unchanged. However, the best overall result is obtained by full translation, R3, with Macro-F1 0.4647 and accuracy 0.4943. R1 improves overall Macro-F1 from 0.3785 to 0.4434, R2 reaches 0.4589, and R3 is strongest.
The per-language gains inside the low-tier on SIB-200 are concentrated in Telugu (+0.4105), Bengali (+0.2914), Amharic (+0.2549), and Swahili (+0.2011), whereas Afrikaans improves only marginally (+0.0135). On MASSIVE, gains are concentrated in Swahili (+0.3116), Telugu (+0.2796), Amharic (+0.2406), Bengali (+0.2297), and Afrikaans (+0.1375).
The paper reports routing effects at the tier level rather than through a single global efficiency score, reporting delta Macro-F1 relative to the multilingual-only baseline and mean latency by resource tier. The paired error analysis shows that R1 low-tier gains reflect many more corrected predictions than newly introduced errors: on SIB-200, 329 fixed versus 84 regressed examples; on MASSIVE, 4606 fixed versus 820 regressed examples, both with p < 0.001.
The authors conclude that the repeated low-tier gain across both benchmarks is the most stable empirical result, while the optimal routing boundary depends on the task. They argue for reporting multilingual deployment outcomes at the point where interventions are applied rather than collapsing all languages into a single system-level score.
Improvements for AI systems
Improvements to AI Systems:
-
Resource-Adaptive Routing for Multilingual Models: Implement a static, language-tier-based routing mechanism that dynamically selects between a direct multilingual inference path and a translation-to-English path, based on per-language performance tiers (high, mid, low). This avoids fine-tuning and reduces computational cost while boosting low-resource language accuracy.
-
Task-Specific Routing Boundary Optimization: Instead of a universal routing policy, allow the system to learn or tune the translation boundary (which tiers to translate) per task or dataset. For topic classification, translate only the lowest tier; for intent classification, translate all tiers for best overall performance. This yields up to +0.22 Macro-F1 gains on low-resource languages without degrading high-resource performance.
-
Cost-Aware Latency Control: Integrate a latency budget into routing decisions. The system can choose a configuration (e.g., R1 vs. R2) that maximizes accuracy while respecting a user-defined mean latency constraint, preventing unnecessary translation overhead (e.g., avoiding R2 when it adds 0.13s latency with no quality gain).
-
Error-Correction Prioritization: Use paired error analysis to guide routing. The system can identify languages where translation yields a high ratio of corrected predictions to newly introduced errors (e.g., 329:84 on SIB-200). This enables proactive routing for languages with high correction potential, while keeping stable languages on the direct path.
-
Self-Hosted Deployment with Lightweight Encoders: Replace large multilingual models with compact sentence encoders (e.g., MiniLM variants) and OPUS-MT/NLLB translation models, enabling fully on-premise, privacy-preserving classification with no external API calls. This reduces infrastructure costs and improves inference speed for short-text tasks.
What the Improved AI System Can Do:
-
Classify short texts in 15+ languages with higher accuracy for low-resource languages (e.g., Telugu, Bengali, Amharic, Swahili) by selectively translating only those languages, achieving gains of +0.20 to +0.41 Macro-F1.
-
Automatically adapt its translation strategy based on the task type (e.g., topic vs. intent classification) to find the optimal accuracy-latency trade-off.
-
Operate entirely offline with minimal computational footprint, making it suitable for edge devices or cost-sensitive production environments.
-
Provide transparent per-tier performance reports (e.g., Macro-F1, latency, error correction counts) rather than a single opaque score, enabling better debugging and user trust.
-
Dynamically avoid unnecessary translations for high-resource languages (e.g., English, Spanish) to keep latency low, while ensuring low-resource languages receive the translation boost they need.
Abstract
Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages. Uniform inference policies are simple to deploy, but they assume that all languages are equally well served. In this work, we evaluate a fixed-list routing strategy that keeps stronger languages on a direct multilingual path and selectively sends weaker languages through translation into English before zero-shot classification. The pipeline is fully self-hosted, uses pretrained compact sentence encoders, and requires no task-specific fine-tuning. We test the approach on two benchmarks chosen to differ in scale and label granularity: a 15-language subset of SIB-200 for seven-way topic classification and a 15-locale subset of MASSIVE for intent classification over an official 60-intent inventory. On SIB-200, the best overall configuration is R1, which translates only the low-resource tier: high-tier and mid-tier Macro-F1 remain unchanged, while low-tier Macro-F1 rises from 0.4632 to 0.6828. On the MASSIVE subset, the same low-tier intervention raises low-tier Macro-F1 from 0.2143 to 0.4417, but the best overall result is obtained by full translation, R3, at Macro-F1 0.4647. Across these two benchmarks, selective translation is a reliable intervention for weaker languages, whereas the optimal routing boundary depends on the task. We therefore report routing through tier-level quality gains and tier-level latency rather than a single global efficiency score.
Sources
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering