Long-Context Demonstration Selection Using State Space Models
cs.LG, cs.CL
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 16 pages
Code: https://github.com/VirtuosoResearch/Long-context-demonstration-selection
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study the problem of demonstration selection, which involves selecting a subset of examples for prepending to a query to a language model.
Terminology
Abstract
We study the problem of demonstration selection, which involves selecting a subset of examples for prepending to a query to a language model. This problem is closely related to in-context learning and language model inference. Since the inference cost of a transformer model scales quadratically with sequence length, the selection problem becomes especially challenging in a long-context scenario. In this paper, we tackle this problem by building on state space models (SSMs), which require only linear inference time given the input. Our approach involves two algorithms. The first learns a small set of SSMs through distillation of a (trained) transformer model. We partition all the layers into consecutive groups. Then for each group, we estimate a separate state space model to replicate the input-output behavior within the adjacent layers. Second, we map the distilled model outputs to a small set of tokens, and apply these embeddings for demonstration selection in downstream applications. We perform extensive experiments in both synthetic and real-world datasets to validate our approach. We demonstrate that the distilled SSMs only incur an approximation error of less than 0.7% relative to the true output. In downstream evaluation, we show that on several text classification and reasoning tasks, our approach reduces FLOPs by 14.2 times and improves accuracy by 6.48% relative to baseline demonstration selection methods.
Sources
- Efficiently Learning Branching Networks for Multitask Algorithmic Reasoning
- Data-Driven Switchback Experiments: Theoretical Tradeoffs and Empirical Bayes Designs
- Scalable Multi-Objective and Meta Reinforcement Learning via Gradient Estimation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks