Selection Is Retrieval, Abstention Is Not: On-Device Tool Routing over 70 Korean-English Actions
cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 12 pages, 4 figures, 13 tables
Code: https://github.com/memovai/mimimodel
License: http://creativecommons.org/licenses/by/4.0/
The gist: An AI assistant that calls tools makes two decisions on every request: which tool to invoke, and whether any available tool applies.
Terminology
Abstract
An AI assistant that calls tools makes two decisions on every request: which tool to invoke, and whether any available tool applies. In the usual design a single language model makes both, by emitting a call or by declining to emit one. On a device that has to answer without a server, the language model is what makes that design expensive, dominating both the latency and the memory of the router. The common alternative is to remove the model completely and rank the catalog of local actions with a retriever instead. That substitution is not symmetric across the two decisions. A retriever returns its highest-scoring candidate for every input and cannot signal that the catalog holds no valid action. Our earlier study found that constraining a decoder to a tool grammar repairs malformed output without improving the choice. What the substitution costs in each decision has not been measured. We evaluate the two decisions separately over 600 Korean and English requests and a catalog of 70 local actions. The router may also ask for a missing slot, reply, or delegate. Half the in-catalog requests reuse catalog vocabulary and half paraphrase it, separating lexical overlap from the action requested. Character 3-gram BM25 selects 162 of 164 lexically matched requests and 85 of 166 paraphrases. Restricting the candidate set to seven raises the paraphrase figure to a mean of 0.825 over five trials. No classifier over its score features separates in-catalog from out-of-catalog above 0.697 area under the curve, where the frozen encoder multilingual-e5-base reaches 0.806. Using that encoder for abstention alone keeps 376 of the requests local and misroutes 9 of the 150 needing delegation. Abstention, not selection, is where a neural component is required. A neural ranker improves every quality metric and is rejected on latency and memory rather than accuracy.
Sources
- Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation
- Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents
- TinyAgent: Function Calling at the Edge
- SLM-Based Agentic AI with P-C-G: Optimized for Korean Tool Use
- Beyond pass@1: A Reliability Science Framework for Long-Horizon LLM Agents
- Most of the LLM Routing Gap Is Task Type
- Repair, Not Improvement: Decomposing Constrained Decoding in Tool-Call Abstention
- FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs
- A Controlled Study of Attention-Only Transformers
- Multilingual E5 Text Embeddings: A Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering