MILO: Efficient Many-shot In-Context Learning with Block-wise Low-rank Compression
cs.CL
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- Many-Shot In-Context Learning
- In-context Learning with Retrieved Demonstrations for Language Models: A Survey
- Language Models are Few-Shot Learners
- ReduNet: A White-box Deep Network from the Principle of Maximizing Rate Reduction
- Towards Compute-Optimal Many-Shot In-Context Learning
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
- Language model compression with weighted low-rank factorization
- LoRA: Low-Rank Adaptation of Large Language Models
- Many-Shot In-Context Learning in Multimodal Foundation Models
- An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction
- MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
- Benchmarking Natural Language Understanding Services for building Conversational Agents
- Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- Breaking the Bank with ChatGPT: Few-Shot Text Classification for Finance
- s1: Simple test-time scaling
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering