Compressing Sequences in the Latent Embedding Space: K-Token Merging for Large Language Models
cs.CL, cs.AI
Submitted: 2026-04-16
Updated: 2026-09-08
Comments: Accepted to EMNLP 2026
Code: https://github.com/shsjxzh/K-Token-Merging
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) incur significant computational and memory costs when processing long prompts, as full self-attention scales quadratically with input length.
Terminology
Abstract
Large Language Models (LLMs) incur significant computational and memory costs when processing long prompts, as full self-attention scales quadratically with input length. Token compression aims to address this challenge by reducing the number of tokens representing inputs. However, existing prompt-compression approaches primarily operate in token space and overlook inefficiencies in the latent embedding space. In this paper, we propose K-Token Merging, a latent-space compression framework that merges each contiguous block of K token embeddings into a single embedding via a lightweight encoder. The compressed sequence is processed by a LoRA-adapted LLM, while generation remains in the original vocabulary. Experiments on structural reasoning (Textualized Tree), sentiment classification (Amazon Reviews), and code editing (CommitPackFT) show that K-Token Merging lies on the Pareto frontier of performance vs. compression, achieving up to 75% input length reduction with minimal performance degradation. Code is available at https://github.com/shsjxzh/K-Token-Merging.
Sources
- Token Merging: Your ViT But Faster
- UniICL: An Efficient Unified Framework Unifying Compression, Selection, and Generation
- In-context Autoencoder for Context Compression in a Large Language Model
- Lossless Token Sequence Compression via Meta-Tokens
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders
- Decoupled Weight Decay Regularization
- LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression
- LLM/Agent-as-Data-Analyst: A Survey
- Scaling Embedding Layers in Language Models
- A Survey of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering