GMTRouter: Personalized LLM Router over Multi-turn User Interactions
cs.CL, cs.LG
Submitted: 2025-10-29
Updated: 2026-09-02
Comments: Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026
Code: https://github.com/ulab-uiuc/GMTRouter
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Large Language Model (LLM) routing has demonstrated strong capability in balancing response quality with computational cost.
Terminology
Abstract
Large Language Model (LLM) routing has demonstrated strong capability in balancing response quality with computational cost. As users exhibit diverse preferences, personalization has attracted increasing attention in LLM routing, since even identical queries may require different models to generate responses tailored to individual needs. However, existing approaches are not fully personalized and often fail to faithfully capture the complex interactions between users and LLMs. Moreover, user preference data is typically scarce and inconsistent in format, which limits the effectiveness of methods that directly leverage user-specific data. To address these challenges, we propose GMTRouter, which represents multi-turn user-LLM interactions as a heterogeneous graph with five node types: user, LLM, query, response and turn, thereby maximally preserving the rich relational structure of the interaction. Through a lightweight inductive graph learning framework combined with a tailored user-conditioned graph sampling mechanism, GMTRouter learns to capture user preferences from few-shot data, enabling effective personalization. Extensive experiments demonstrate that GMTRouter outperforms the strongest baselines, achieving up to a 0.108 absolute improvement in accuracy and a 0.124 improvement in AUC. More importantly, we further demonstrate that GMTRouter can adapt to new users using only few-shot data, without extensive fine-tuning. The code for GMTRouter is publicly available at https://github.com/ulab-uiuc/GMTRouter.
Sources
- GPT-4 Technical Report
- Personalized Graph-Based Retrieval for Large Language Models
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- You are AllSet: A Multiset Function Framework for Hypergraph Neural Networks
- Training Verifiers to Solve Math Word Problems
- Cost-Effective Online Multi-LLM Selection with Versatile Reward Models
- A Unified Approach to Routing and Cascading for LLMs
- Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
- A Survey on In-context Learning
- GraphRouter: A Graph-based Router for LLM Selections
- FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
- Aligning LLM Agents by Learning Latent Preference from User Edits
- The Llama 3 Herd of Models
- Inductive Representation Learning on Large Graphs
- RouterBench: A Benchmark for Multi-LLM Routing System
- Unsupervised Dense Information Retrieval with Contrastive Learning
- Mixtral of Experts
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering