When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use

arXiv:2608.11715 · cs.CL, cs.AI · Submitted 2026-08-12 · Read on arXiv

Siddharth Chauhan, Thomas Butler, Abhishek Singhania, Pankaj Porwal, Honey Gupta

Amazon

cs.CL, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: This paper investigates a failure mode in multilingual API calling called Argument Language Mismatch (ALM), where a model selects the correct tool but generates argument values in a language

Terminology

Summary

This paper investigates a failure mode in multilingual API calling called Argument Language Mismatch (ALM), where a model selects the correct tool but generates argument values in a language inconsistent with the user input. The authors revisit post-training strategies to mitigate ALM, finding that supervised fine-tuning (SFT) provides a strong baseline, substantially improving argument language consistency and end-to-end function call accuracy. They then examine whether reinforcement learning (RL) with structured, argument-aware rewards offers additional benefits, finding that methods like Group Relative Policy Optimization (GRPO) can improve language consistency and better preserve general reasoning ability, but these gains are incremental and most pronounced in generalization and multi-objective trade-offs.

The paper makes three key contributions: formalizing ALM as a distinct failure mode, showing that SFT provides a strong baseline for mitigating ALM, and evaluating RL with structured rewards to find incremental gains over SFT. To support the study, the authors construct a multilingual extension of the Berkeley Function Calling benchmark covering Spanish, French, Italian, and Dutch, with two splits: Split-1 (learnability) with moderate API overlap and Split-2 (generalization) with minimal API overlap. Models are trained on Spanish only and evaluated on Spanish and unseen languages.

The evaluation uses a hierarchical set of metrics: Tool Invocation Detection (TID), Tool Selection Accuracy (TSA), Argument Completion Accuracy (ACA), Argument Language Consistency (ALC), and Function Call Match (FCM), forming a strict hierarchy FCM ≤ ALC ≤ ACA ≤ TSA ≤ TID. The primary objective is to maximize ALC without sacrificing FCM or general reasoning ability.

For training, the paper compares SFT with PPO and GRPO. The reward design progresses from RM-1 (sparse binary reward) to RM-2 (hierarchical step reward) to RM-3 (argument-factorized reward), with RM-3 providing continuous feedback per argument value. Token-level reward weighting is also introduced, upweighting argument-value tokens with β ∈ 1.5, 3.

Key results show that SFT substantially improves ALC and FCM over the base model, resolving a large fraction of ALM errors. When comparing at the best validation checkpoint, SFT achieves 79.1 ALC / 67.4 FCM on Split-1, exceeding GRPO (74.0 / 55.3) and SFT+GRPO (79.3 / 61.3) on FCM. However, RL methods, particularly GRPO, better preserve English reasoning ability: the validation-selected SFT model drops from 70.8 to 62.2 on English MGSM (−8.6 points), while GRPO leaves English reasoning essentially intact (70.4).

The reward model ablation shows monotonic improvement from RM-1 to RM-3: RM-1 achieves 61.3 ALC / 43.3 FCM, RM-2 achieves 72.2 / 51.0, and RM-3 achieves 74.0 / 55.3. GRPO consistently outperforms PPO under identical reward formulations, with GRPO reaching 81.2 ALC / 66.9 FCM versus PPO's 72.6 / 58.4. Token-level reward weighting improves GRPO (β=3 gives 77.74 ALC) but destabilizes PPO (β=3 drops to 50.81 ALC).

Cross-lingual transfer results show GRPO provides consistent improvements across unseen languages (Italian, Dutch, French), while SFT shows negative transfer on Dutch (−1.89 points). Model scaling experiments show smaller models trained with GRPO often match or exceed larger models trained with SFT: 7B GRPO achieves 68.10 ALC versus 7B SFT's 65.14, and 32B GRPO reaches 73.78 ALC versus 32B SFT's 67.59.

The paper concludes that much of the multilingual API grounding performance can be achieved through careful supervised training and model selection, with RL providing targeted rather than fundamental improvements. The discussion notes that ALM is primarily a surface-level alignment issue learnable through imitation, while RL helps most in generalization settings and multi-objective trade-offs. The authors emphasize the importance of strong baselines and controlled comparisons when evaluating post-training strategies for structured generation.

Improvements for AI systems

Improvements to AI Systems:

  1. Add a language-consistency verification layer for tool-call arguments. The system can detect when a model selects the correct API but generates argument values in a different language than the user’s query, then automatically re-generate those arguments in the correct language before execution.

  2. Implement a hierarchical reward function for training on structured outputs. Instead of a single binary reward for a full function call, the system can provide separate, continuous rewards for tool selection, argument completion, and argument language consistency—enabling more granular learning and reducing failure cascades.

  3. Use token-level reward weighting during reinforcement learning. The system can upweight loss on argument-value tokens (e.g., β=3) during GRPO training, improving language consistency without destabilizing training, as opposed to uniform token weighting.

  4. Prefer GRPO over PPO for multilingual API fine-tuning. The system can adopt GRPO as the default RL algorithm for structured generation tasks, as it yields higher argument language consistency (81.2 vs. 72.6 ALC) and better preserves general reasoning ability under identical reward designs.

  5. Adopt a two-stage training pipeline: SFT first, then GRPO with argument-factorized rewards. The system can use SFT to establish a strong baseline for tool-calling accuracy, then apply GRPO with RM-3 (per-argument continuous feedback) to improve cross-lingual generalization and multi-objective trade-offs without sacrificing function call match.

  6. Add a cross-lingual negative-transfer detector. The system can monitor performance on unseen languages during fine-tuning and flag when SFT degrades performance (e.g., Dutch), then automatically switch to GRPO or adjust training data to prevent such degradation.

  7. Enable model-scaling efficiency via RL. The system can train smaller models (e.g., 7B) with GRPO to match or exceed the multilingual API performance of larger models (e.g., 32B) trained with SFT alone, reducing computational cost while maintaining quality.

  8. Integrate a multi-objective optimization module. The system can balance argument language consistency (ALC) with function call match (FCM) and general reasoning (e.g., English MGSM) by using GRPO’s inherent trade-off handling, avoiding the reasoning degradation seen with SFT (−8.6 points).

What the Improved AI System Can Do:

  • Execute API calls in multilingual settings (Spanish, French, Italian, Dutch) with high argument language consistency, even for unseen languages, without sacrificing tool selection accuracy.

  • Automatically correct or flag language mismatches in generated arguments before execution, reducing runtime errors.

  • Train more efficiently on structured generation tasks, using smaller models that outperform larger SFT-only models.

  • Preserve general reasoning ability (e.g., math problems) while improving multilingual tool-calling, enabling deployment in mixed-task assistants.

  • Provide explainable reward signals during training, allowing developers to pinpoint whether failures stem from tool selection, argument completion, or language inconsistency.

  • Generalize to new languages with minimal negative transfer, thanks to GRPO’s robustness over SFT in cross-lingual settings.

Abstract

The reliability of Large Language Models (LLMs) for API calling degrades in multilingual settings. A common failure occurs when a model selects the correct tool but generates argument values in an inconsistent language, which we term Argument Language Mismatch (ALM). Although semantically correct, such outputs are operationally invalid and not captured by standard API-calling metrics. We revisit post-training strategies for mitigating ALM and find that, in our benchmark, supervised fine-tuning (SFT) provides a strong baseline, substantially improving argument language consistency and end-to-end function call accuracy. Under consistent model selection, SFT achieves performance comparable to, and sometimes exceeding more complex reinforcement learning (RL) approaches. We further examine whether RL with structured, argument-aware rewards offers additional benefits. While methods such as Group Relative Policy Optimization (GRPO) can improve language consistency and better preserve general reasoning ability, these gains are incremental and most pronounced in generalization and multi-objective trade-offs. Overall, our results suggest that much of the performance in multilingual API grounding can be achieved through careful supervised training, with RL providing targeted rather than fundamental improvements.

Sources

Related papers