Distilling Sequential Computation in Transformer Language Models
cs.CL
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/Zesearch/Umim-LLM
Project page: http://skylion007.github.io/OpenWebTextCorpus
Terminology
Sources
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
- ROUGE 2.0: Updated and Improved Measures for Evaluation of Summarization Tasks
- zip2zip: Inference-Time Adaptive Tokenization via Online Compression
- The Llama 3 Herd of Models
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
- MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
- Copy Is All You Need
- Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference
- Unlocking Context Constraints of LLMs: Enhancing Context Efficiency of LLMs with Self-Information-Based Content Filtering
- Compressing Context to Enhance Inference Efficiency of Large Language Models
- 500xCompressor: Generalized Prompt Compression for Large Language Models
- Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference
- SuperBPE: Space Travel for Language Models
- Pointer Sentinel Mixture Models
- Byte Latent Transformer: Patches Scale Better Than Tokens
- LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression
- Flexibly Scaling Large Language Models Contexts Through Extensible Tokenization
- ProCut: LLM Prompt Compression via Attribution Estimation
- Cmprsr: Abstractive Token-Level Question-Agnostic Prompt Compressor
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering