UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification
cs.CL
Submitted: 2026-05-07
Updated: 2026-09-25
Code: https://github.com/qhfan/UniPrefill
Terminology
Sources
- Qwen Technical Report
- VSPrefill: Vertical-Slash Sparse Attention with Lightweight Indexing for Long-Context Prefilling
- The Llama 3 Herd of Models
- FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
- LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Mistral 7B
- Gemma 3 Technical Report
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- SnapKV: LLM Knows What You are Looking for Before Generation
- MoBA: Mixture of Block Attention for Long-Context LLMs
- Retentive Network: A Successor to Transformer for Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- ProxyAttn: Guided Sparse Attention via Representative Heads
- MiMo-V2-Flash Technical Report
- Optimizing Mixture of Block Attention
- Qwen3 Technical Report
- Qwen2.5 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering