Repair, Not Improvement: Decomposing Constrained Decoding in Tool-Call Abstention
cs.CL
Submitted: 2026-08-14
Updated: 2026-09-05
Comments: 24 pages, 4 figures, 21 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: A tool-calling router has to pick the right tool when one applies and decline when none does.
Terminology
Abstract
A tool-calling router has to pick the right tool when one applies and decline when none does. Restricting the decoder to a grammar over the tool names is the standard remedy for the first, and on small models it buys a large accuracy gain. Recent work separates the loss caused by asking for a format from the loss caused by enforcing it at decode time. The second is small, which made enforcement look nearly free. The same work declines to extend that to function calling, where a constraint decides which answers exist rather than how one is written. Declining to call anything is the answer it most easily removes, and the one a router can least afford to lose. A grammar decides which tokens may be emitted and where generation stops, and a two-condition design charges both to the restriction. We therefore run three conditions over one prompt: free generation, generation stopped at the first line, and both applied together. We evaluate open-weight models from 0.6B to 4B on the same items in English and Korean, comparing the languages item by item. The two-condition contrast is negative on abstention accuracy in four of six cells with intervals excluding zero, and positive in none, costing-29.5 points at worst. On the smallest model in Korean the stop costs-20.0 points, the restriction returns +19.5, and together they leave-0.5. What the restriction gives back is readable output, not judgment. Of the 698 abstentions it repairs, 545 had no readable answer at all and 0 were correct decisions the scoring rule rejected. On items that do need a tool the contrast is positive throughout, and abstention is reported first because it is the registered measure. Both preregistered claims about language fail: Korean does not lose more of the abstentions it holds without the constraint, and the removed mass does not explain what does.
Sources
- Mitigating Bias in Locally Constrained Decoding via Tractable Proposals
- Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents
- Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints
- Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo
- (G)I-DLE: Generative Inference via Distribution-preserving Logit Exclusion with KL Divergence Minimization for Constrained Decoding
- Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
- The Format Tax
- Grammar-Aligned Decoding
- The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models
- Preregistering NLP Research
- From Hallucination to Structure Snowballing: The Alignment Tax of Constrained Decoding in LLM Reflection
- The Hidden Cost of Structured Generation in LLMs: Draft-Conditioned Constrained Decoding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering