Investigating Knowledge Transfer Across Interactive Dialogue Games
cs.CL
Submitted: 2026-08-25
Updated: 2026-08-25
License: http://creativecommons.org/licenses/by/4.0/
The gist: Dialogue games represent a challenging setting where complex cognitive skills are required to accomplish tasks while coordinating with other players.
Terminology
Abstract
Dialogue games represent a challenging setting where complex cognitive skills are required to accomplish tasks while coordinating with other players. Considering that language represents an interface for both understanding the game rules and executing actions, it is reasonable to assume that training on a specific language game will enhance specific capabilities that might be relevant for other tasks as well. Motivated by this rationale, in this paper, we investigate how knowledge transfers across different dialogue games. We study transferability by finetuning LLM models on games from the clembench suite (Chalamalasetti et al., 2023) and performing two analyses: i) we derive a task-transferability graph using a binary integer optimization program from Zamir et al. (2018), using task performance as the main metric; and ii) we compute task vectors (Ilharco et al., 2022) for each game to study similarities across finetuned models and their task transferability. In our first analysis, we find that some games benefit more from transfer than finetuning, and that the visuospatial family (e.g., exploration games) transfers best. With our task vector analysis instead, we find that similarity-based approaches capture game-role relationships but almost no transferability patterns, suggesting that more complex metrics are required.
Sources
- Measuring Massive Multitask Language Understanding
- Editing Models with Task Arithmetic
- Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
- GameEval: Evaluating LLMs on Conversational Games
- Dialogue Games for Benchmarking Language Understanding: Motivation, Taxonomy, Strategy
- GLU Variants Improve Transformer
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering