NOSA: Native and Offloadable Sparse Attention
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "NOSA: Native and Offloadable Sparse Attention".
Tom: Decoding throughput improvements from larger inference batches are limited by GPU memory, which is largely consumed by the key-value (KV) cache.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So we’ve been deep into the mechanics of NOSA, and now we need to talk about what this whole thing means for the language model world regarding its title and authors.
Jane: I think focusing on the title, "NOSA: Native and Offloadable Sparse Attention," really frames how clever these authors were in solving that memory puzzle for us.
Lu: Exactly! The authors managed to bake this sparsity constraint right into their training process, which is a smart way to handle the long-term efficiency issues we’ve seen with KV caches.
Meng: From an engineering standpoint, the "Native" part suggests they built something integrated for offloading rather than just slapping on a patch later. That makes implementing it cleaner on production systems.
Lalam: For me, this work speaks to a future where we can train and run much larger models without immediately needing massive amounts of dedicated GPU memory, which opens up so many possibilities for diverse applications.
Tom: It really highlights that the authors weren't just optimizing a component; they designed an entire system meant to handle the real-world constraints of long contexts.
Jane: And their ability to maintain performance on those long tasks is what’s most compelling, showing that we don't have to sacrifice quality just for speed.
Lu: The core implication is that we can push the boundaries of context length and model size because the bottleneck shifts from simple memory limitations to something more manageable, like carefully controlled attention locality.
Meng: I see the practical impact here as making deployment much more flexible; if we can reliably offload the cache like they did, we don't have to restrict ourselves to smaller models just because of hardware constraints.
Lalam: It fundamentally improves how we interact with complex information across vast sequences, fostering a new culture where deep reasoning over massive data is much more accessible.
Tom: It’s clear that this paper lays down a solid foundation for how we approach sparse attention in future architectures, and that sets the stage for some really exciting follow-up research into scaling these techniques further.
Conclusion: Tom: So, to wrap up this discussion on NOSA: Native and Offloadable Sparse Attention, we’ve seen how the authors managed to build a smart system that directly addresses the memory challenges of long context processing by training in a way that favors efficiency.
Jane: I think focusing on the title itself really tells us a lot about their approach; "Native" suggests they didn't just add an afterthought, but built this sparsity constraint right into the core training process, which is a big deal for robustness.
Lu: Exactly! They formalized that locality requirement during training by splitting the selection process into query-aware and query-agnostic parts, showing a deep understanding of how attention needs to be managed over many decoding steps.
Meng: From an engineering standpoint, the "Offloadable" aspect is crucial because it means this isn't just a theoretical improvement; it’s designed to work with actual deployment systems where memory management is tight.
Lalam: For me, this paper suggests that we can start thinking about models that can handle massive amounts of data without immediately hitting a hard hardware wall, opening up incredible avenues for complex applications.
Tom: It really highlights that the authors didn't just fix a single problem; they designed an entire framework meant to handle the real-world constraints of scaling language models effectively.
Jane: And their ability to keep performance strong on those long tasks while boosting throughput is what makes this research so compelling for everyone in the community.
Lu: The core implication here is that we can push the boundaries of context length and model size because the bottleneck shifts from simple memory limitations to something more manageable, like carefully controlled attention locality.
Meng: I see a huge practical impact here as making deployment much more flexible; if we can reliably offload the cache like they did, we don't have to restrict ourselves to smaller models just because of hardware constraints.
Lalam: It fundamentally improves how we interact with complex information across vast sequences, fostering a new culture where deep reasoning over massive data is much more accessible for everyone.
Tom: It’s clear that this paper lays down a solid foundation for how we approach sparse attention in future architectures, and it sets the stage for some really exciting follow-up research into scaling these techniques further.
Jane: And that leads us perfectly into looking at how these results translate into real-world performance metrics across different model sizes.
NLP Group at Tsinghua University · OpenBMB (China) · BUPT (Beijing) · School of Computer Science and Technology at Beijing Institute of Technology
cs.CL, cs.AI, cs.LG
Submitted: 2025-10-15
Updated: 2026-09-28
Code: https://github.com/thunlp/NOSA
Importance score: 85/100
The gist: Decoding throughput improvements from larger inference batches are limited by GPU memory, which is largely consumed by the key-value (KV) cache.
Key concepts
- KV Cache Offloading
- This is the problem where the Key-Value cache (which stores previously processed information) consumes too much GPU memory, limiting how much text a model can process at once. NOSA aims to reduce this memory pressure by intelligently selecting which parts of the cache to keep and which ones to offload.
- Locality Constraint ($orall t o n, ar{ ext{t}} o t' > t$)
- This is the core training innovation. It forces the model to select tokens for different decoding steps that are close together in sequence. This constraint ensures that once a token is chosen, it is likely to be needed again soon, improving data reuse and speeding up generation.
- Query-Aware Selection
- This part of the mechanism selects tokens based on their relevance to the current query (the token being generated). It preserves model performance by not adding extra restrictions. It ensures that important information is always prioritized during the selection process.
Terminology
Summary
Decoding throughput improvements from larger inference batches are limited by GPU memory, which is largely consumed by the key-value (KV) cache. This paper proposes NOSA, a trainable sparse attention mechanism natively designed for KV cache offloading and an accompanying inference system NOSI, to achieve high decoding throughput with minimal performance degradation.
The gist
NOSA outperforms KV cache offloading baselines on general, long-context input, and long-generation tasks while boosting decoding throughput by up to 5.04×, 1.92×, and 1.83× over FullAttn, InfLLMv2, and ShadowKV.
How it works
The core innovation of NOSA is adding a locality constraint during training to improve KV selection locality across decoding steps, formalized by imposing a lower bound on the overlapping rate of selected tokens between adjacent decoding steps: ∀t ∈ 2, · · ·, n, γ(t) ≥ γ0
(Equation 3). This constraint is enforced by partitioning the selected tokens into two subsets: query-aware selection and query-agnostic selection. The query-agnostic selection is implemented using an eviction head, which assigns importance scores to each KV pair. For the query-aware selection, a Top-k function is applied to the scores of the selected tokens, while for the query-agnostic selection, a Top-k function is applied to these importance scores.
Key Design Components
-
The mechanism decomposes KV selection into
query-aware and query-agnostic components.
The query-aware component preserves model performance by not adding constraints, while the query-agnostic component imposes an eviction constraint: "if a token is not selected at decoding step t, it must not be selected at any subsequent step t′ > t." -
The system adopts a two-stage, block-wise selection strategy on top of InfLLMv2. This involves mean pooling and max pooling to derive selection scores for query-aware and query-agnostic selections across blocks.
-
The importance score for the eviction head is computed using a trainable mechanism:
s e j = τ (v jW1)W2
(Equation 4), where τ is implemented as the softplus function in DMA’s design, and W1, W2 are trainable parameters. -
The selection process follows Equation 7, which selects up to k/nb blocks by prioritizing query-aware selection based on a budget kq and filling the remainder with query-agnostic blocks based on a budget ke.
System Optimization (NOSI)
To fully unlock NOSA’s efficiency, the paper introduces NOSI, an inference system specifically tailored for NOSA. Key optimizations include:
-
Kernel Fusion: Utilizing FlashInfer implementations for LayerNorm, RoPE, and FFN, and fusing the eviction heads with splitting QKV tensor into a single kernel.
-
Memory Layout and Block Selection: Organizing the KV cache into minimal block-sized units on both GPU and CPU, only swapping out blocks not used in the current step.
-
Offloading Communication: Implementing a custom Triton kernel for host-to-device transfers that leverages Unified Virtual Addressing (UVA) to access CPU memory directly, achieving up to 83% of peak PCIe bandwidth compared to vanilla PyTorch implementations.
Empirical Results and Analysis
The paper demonstrates the effectiveness of NOSA through extensive experiments on 1B, 3B, and 8B models across various benchmarks (LongBench, HELMET).
-
Performance: NOSA
outperforms KV cache offloading baselines on general, long-context input, and long-generation tasks.
It maintains consistent performance where other methods suffer fromtraining–inference mismatch,
especially in long-generation reasoning tasks where training-free offloading (ShadowKV) shows asharply increasing perplexity.
-
Throughput: NOSA achieves significant speedups, boosting decoding throughput by up to 5.04× on long-context input (96K) and 1.83× over ShadowKV on general tasks, showing that
Higher locality leads to faster decoding speed
as NOSA achieves better locality than InfLLMv2. -
Tunability: The ratio kq/k, which controls the minimum locality γ0 = ke/k, allows for a trade-off:
Decreasing kq/k enforces higher locality, which increases decoding throughput but weakens query-aware selection, limiting recall and reducing task performance.
The authors set kq/k = 0.25 in main experiments to balance these factors. -
System Necessity: The wall-time breakdown analysis shows that NOSI exposes the
true communication overhead from other performance bottlenecks,
proving that optimizing KV-loading time yields a "meaningful evaluation of KV-loading optimizations.
Improvements for AI systems
Here are the specific improvements and capabilities for an AI system based on the NOSA (Native and Offloadable Sparse Attention) paper:
)1. Enhanced Decoding Throughput via Native Locality Constraint:
The system will implement a native sparse attention mechanism that explicitly enforces a minimum selection locality threshold, formalized as a lower bound on the overlap of selected tokens between consecutive decoding steps, i.e., ensuring the fraction of selected tokens from step-t is greater than or equal to a predefined level, γ(t) ≥ γ0.
)2. Optimized KV Cache Offloading via Dual Selection Strategy:
The system will decompose KV cache selection into two components:
-
A query-aware selection component (to maintain long-context recall and performance).
-
A query-agnostic selection component (using a trainable eviction head/DMA variant to bound the number of CPU–GPU transfers and enforce an eviction policy, preventing catastrophic training–inference mismatch).
)3. Efficient Host–Device Communication via Block-wise Selection:
The system will utilize a block-wise selection strategy over element-wise selection, combining mean pooling (within blocks) and max pooling (across blocks). This approach is specifically designed to optimize PCIe bandwidth utilization by ensuring that KV cache transfers are consolidated into larger, more efficient block movements rather than fragmented element transfers.
)4. System-Level Inference Optimization via NOSI:
The inference system will be optimized using NOSI, which includes:
-
Kernel Fusion of computationally expensive components (like eviction heads and pooling stages) into single Triton kernels to reduce kernel launch overheads.
-
The use of a custom CUDA kernel leveraging Unified Virtual Addressing (UVA) for direct, fine-grained host–to–device access, achieving up to 83% of peak PCIe bandwidth for KV cache transfers.
)5. Improved Long-Context Performance and Robustness:
The system will be trained using the NOSA mechanism starting from long-context continual pretraining (using methods like InfLLMv2-data-5B). This training approach ensures that the sparse patterns learned during training are consistent during inference, preventing the error accumulation that degrades quality in training-free offloading methods (like ShadowKV) during long generations.
)6. Tunable Performance/Efficiency Trade-off:
The system will incorporate a hyperparameter, the ratio of query-aware to query-agnostic selection budget (kq/k), allowing for dynamic adjustment at inference time. This enables the user to trade off between maximizing long-context performance (higher kq/k) and maximizing decoding throughput (lower kq/k).
This improved AI system will be capable of performing the following:
-
Generate significantly longer, high-quality sequences (e.g., 8K or 16K tokens) with minimal drop in perplexity, especially on reasoning tasks, by maintaining consistency between training and inference sparsity patterns.
-
Process extremely long context inputs (e.g., 32K or 64K tokens) on memory-constrained hardware (like consumer GPUs), where it can maintain competitive performance against state-of-the-art baselines that suffer from catastrophic training–inference mismatches.
-
Achieve substantially higher decoding throughput (up to 5.04× over FullAttn) during inference by minimizing the communication overhead associated with offloading the KV cache, making it suitable for high-throughput serving environments.
-
Be highly robust to varying attention sparsity levels (budget k), maintaining strong performance across a wide range of sparsity configurations without significant degradation in recall or reasoning tasks.
Sources
- Program Synthesis with Large Language Models
- Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
- Sparse Attention across Multiple-context KV Cache
- Evaluating Large Language Models Trained on Code
- ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
- Training Verifiers to Solve Math Word Problems
- The Llama 3 Herd of Models
- Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads
- GPT-4o System Card
- PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- MiniCPM4: Ultra-Efficient LLMs on End Devices
- Trainable Dynamic Mask Sparse Attention
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Long-Context Generalization with Sparse Attention
- LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
- FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
- Post-Training Sparse Attention with Double Sparsity
- Hierarchical Sparse Attention Framework for Computationally Efficient Classification of Biological Cells
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering