ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference".
Jane: DeltaKV introduces a residual-based KV cache compression framework that leverages long-range inter-token similarity to significantly reduce memory footprint while maintaining near-lossless performance for long-context Large Language Models.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So Jane and I have been looking at the paper "ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference," and it seems like they’ve tackled a really fundamental problem with how much memory those long context windows are eating.
Jane: Exactly, Tom. It's about the linear growth of KV cache memory becoming a massive bottleneck for deploying powerful LLMs in things like autonomous agents and complex reasoning tasks, which is exactly what this paper addresses.
Lu: I think the core idea they propose is using a residual-based approach to compress the KV cache by encoding token-specific deviations relative to a small set of globally retrieved historical references. That sounds incredibly creative because it moves away from just throwing tokens out when memory gets tight, which is what many other methods do.
Meng: From an engineering standpoint, that reference mechanism sounds complex, so I’m curious how they actually manage that retrieval process without slowing down the inference speed too much.
Lalam: If we can cut the memory footprint by reducing it to twenty-nine percent of the original size while keeping near-lossless accuracy, that has huge implications for how much context our models can handle efficiently in production environments.
Tom: That's what really stands out—reducing that cache memory by a factor of nearly four without sacrificing quality, which is a big deal for deployment on less powerful hardware.
Jane: And they motivate this method by pointing out two key empirical findings: first, that tokens with long-range similarity are actually distributed globally across the context, not just locally near each other.
Lu: That finding about global redundancy is fascinating; it means we have to look beyond just immediate neighbors when deciding what to keep in our cache.
Meng: So, if similarity isn't local, how does the framework manage that long-range retrieval efficiently during a fast inference pass?
Tom: Well, they describe a Strided Reference Selection process where they use a fixed interval stride 's' to maintain a reference set T and then retrieve the top-k nearest tokens from that set based on their L2 distance to calculate the mean reference representation, KV R.
Jane: That sounds like a structured way to find those important historical anchors, and then they use that mean representation as a baseline for calculating what's truly new or residual for each token.
Lalam: I think the mechanism of reconstructing the final KV cache as KV i = KV i + KV R makes perfect sense; it keeps the main reference information intact while only storing the smaller residual parts.
Tom: And they trained this using a hybrid objective function that mixes reconstruction loss with a next-token prediction loss to make sure the compression doesn't destroy the model's ability to actually generate text correctly.
Title and authors: Lu: Combining those two losses ensures that we aren't just compressing data randomly but are preserving both the structural information and the semantic meaning needed for generation.
Meng: I wonder about that training objective, because if you get it wrong in either part, will you end up with a compressed cache that is both inaccurate and unusable?
Jane: That’s exactly what they are trying to manage; they want fidelity without losing the generative capability of the LLM during the compression process.
Tom: Moving on to how they improve upon previous methods, one major point is addressing local similarity bias, which we saw in other works like CacheGen or Chelsea where they only looked at nearby tokens.
Lu: They show that their approach overcomes this by leveraging global similarity across distant tokens; specifically, over sixty percent of the most similar tokens occur at distances greater than sixteen.
Jane: So, instead of relying on local assumptions, DeltaKV uses this global view to find better references for the residuals.
Meng: That makes sense, but they also noted that some existing pipeline approaches introduce irregular memory access patterns and GPU underutilization because of those multi-stage compression methods.
Tom: Right. They are specifically working to be more GPU-friendly by introducing Sparse-vLLM, an inference engine with decoupled memory management and kernels tailored for these sparse and irregular KV layouts.
Lalam: When you combine the compression framework with an inference engine designed for sparsity, it allows them to selectively decompress only the KV pairs required by the sparse attention mask during computation.
Jane: That selective decompression is key because it avoids unnecessary memory I/O when working with sparse attention mechanisms like Sparse-vLLM.
Lu: The architecture involves partitioning the cache into a small set of uncompressed reference tokens and a larger set of compressed tokens, which sets up the entire residual encoding scheme.
Tom: And they're optimizing the kernel execution by using custom Triton kernels, including a Batch L2 Distance kernel for fast reference searching and a Fused Reconstruction kernel that combines everything to minimize GPU memory bandwidth consumption.
Meng: That sounds like a very deep optimization; minimizing that bandwidth usage is critical when dealing with the sheer volume of data in long context scenarios.
Jane: It shows they aren't just focusing on the mathematical compression but on making the actual hardware execution as efficient as possible.
Tom: So, to wrap up this discussion on "ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference," we’ve seen how they use long-range similarity and residual encoding to reduce memory usage significantly while maintaining accuracy.
Title and authors: Jane: The main implication is that we can finally deploy models with extremely long contexts, like those needed for complex reasoning, on consumer hardware without running out of GPU memory.
Lu: It suggests a new direction for KV cache optimization should be focusing on similarity-aware representations rather than just simple eviction or low-rank projections.
Meng: Practically speaking, this means we can serve much larger models with better throughput improvements when integrated with sparse attention methods, which is what the paper shows up to two times higher throughput over vLLM.
Lalam: For our culture, this means we can build systems that handle massive amounts of context reliably and efficiently, opening up new possibilities for applications like high-fidelity document analysis.
Tom: So we’ve seen how they achieve a memory reduction to twenty-nine percent while maintaining near-lossless accuracy through this residual-based KV cache compression framework called DeltaKV.
Jane: It’s a really principled path toward scalable long-context LLM deployment by exploiting global redundancy in KV representations instead of just local similarity assumptions.
Lu: The future work they mention focuses on reducing runtime overhead by moving view and slot management onto GPU kernels and exploring full-pipeline quantization, which could theoretically reduce the footprint to approximately seven point two percent of the original size.
Meng: That goal of reaching that seven point two percent footprint with full-pipeline quantization is ambitious, but if they can do it while keeping it usable, that would be a massive win for energy efficiency and bandwidth usage.
Tom: It really is a lot of work to get there, but the path they’ve laid out using DeltaKV provides a solid foundation for making long-context inference practical.
Jane: Overall, this paper shows us how focusing on semantic residuals relative to retrieved references can lead to substantial memory savings while ensuring that the LLM’s performance remains high across demanding benchmarks.
Lu: The implications for multimodal applications are huge because it tackles the core issue of managing large representations effectively across different modalities when context is long.
Meng: I think this work directly impacts how we architect inference servers, forcing us to rethink how memory is managed at the token level rather than just the page level.
Lalam: For our internal systems, this means we can aim for much higher quality and longer responses in complex reasoning tasks because the underlying infrastructure won't choke on cache size.
Tom: So that’s our wrap-up on "ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference," a paper that offers a very concrete method for tackling KV cache growth through residual encoding and integration with Sparse-vLLM.
Jane: It gives us a lot to think about as we look toward deploying the next generation of large models in demanding, real-world scenarios.
The paper's summary: Tom: So, we’ve got the core summary of ResidualKV laid out: they're using residual encoding based on long-range similarity to slash KV cache memory down to about twenty-nine percent of its original size without losing much accuracy for long contexts.
Jane: That’s a powerful summary because it highlights that their main trick isn't just throwing away old data; it’s intelligently summarizing the important parts into these residuals relative to a few key historical reference points.
Lu: I think what makes this interesting is that the paper explicitly connects the compression strategy to empirical findings showing that semantic similarity happens globally, not just locally, which justifies why they need this long-range retrieval method.
Meng: From an engineering standpoint, that twenty-nine percent reduction is significant; for us running inference servers, knowing we can fit much larger models on the same hardware is huge.
Lalam: For our internal LLM vision, this means we can support significantly longer reasoning chains and document processing tasks natively because the memory constraint that usually kills context length has been substantially loosened.
Tom: Exactly! It's not just about fitting things in; it’s about doing complex reasoning with high-fidelity context for much longer periods.
Jane: And they tied this compression into an inference engine called Sparse-vLLM, which lets the system selectively decompress only what’s needed for a specific attention step, which really makes it practical.
Lu: The integration with Sparse-vLLM is what elevates this beyond just a theoretical compression method; it shows how the mathematical concept of residuals can actually be operationalized in a high-throughput inference environment.
Meng: I’m interested in the performance numbers they showed, specifically that they saw up to two times higher throughput when using this combination with Sparse-vLLM compared to standard vLLM setups in long context scenarios.
Tom: That's the real kicker for deployment; a memory saving is one thing, but doubling the speed makes it usable in demanding applications like autonomous agents, which is what we’re targeting.
Jane: It really moves us toward a future where powerful LLMs aren't just experimental tools but robust infrastructure that can handle truly massive workloads efficiently.
Lu: If we think about the broader AI landscape, this suggests that focusing on similarity-aware representations across long distances is the way forward for scaling context in these models.
Meng: It forces us to reconsider how we manage memory at the token level instead of just managing it at the page level, which is a much more granular and efficient approach for serving models.
Lalam: For our culture, this means we can build systems that handle massive amounts of context reliably and efficiently, opening up new possibilities for high-fidelity document analysis across our entire platform.
Tom: So we’ve seen how DeltaKV uses residual encoding based on global similarity to achieve a twenty-nine percent memory reduction and a potential two times throughput boost when paired with Sparse-vLLM.
Jane: It really shows that by looking at the global structure of the cache, we can find intelligent ways to compress it while preserving what makes those long contexts useful for complex tasks.
Lu: The paper’s main contribution is showing a principled way to exploit this global redundancy rather than relying on local heuristics, which is a major conceptual leap in KV optimization.
Meng: We need to keep an eye on how they handle that training objective combining reconstruction loss and next-token prediction; that balance between accuracy and generation quality is crucial for real-world deployment.
Tom: And we'll be keeping a close watch on those future work directions, especially the goal of reaching even smaller footprints with full-pipeline quantization down to seven point two percent.
The paper's improvements: Tom: So, moving past just the compression itself, the paper lays out several improvements they suggest to make this framework even better for real-world use, focusing on making it faster and more integrated into existing systems.
Jane: That’s right; they aren't stopping at just saving memory; they’re thinking about how to optimize the entire pipeline so that this compression works seamlessly during inference.
Lu: They propose an asymmetric compressor and decompressor design, using things like a SwiGLU block for compression and bias-free linear projection for decompression, which should significantly cut down on the computational overhead during those repeated decoding steps.
Meng: From a practical standpoint, that kind of specialized kernel design is what we need to see; if they can reduce the overhead of reconstruction during every single token generation step, it makes the whole system much more viable for high-speed serving.
Tom: And they aren't just focusing on the core math; they’re pushing toward a full-pipeline quantization strategy, which could theoretically shrink the memory footprint even further down to around seven point two percent of the original size.
Jane: That level of optimization is what makes me really excited because it suggests that we might be able to run these models on much more constrained hardware than we currently think is possible.
Lu: That reduction target would be incredible if achievable, but they also noted a limitation: the method's success heavily depends on identifying those long-range similarities effectively during the reference selection phase, which is where the initial complexity lies.
Meng: I agree with Lu; the effectiveness hinges on that retrieval mechanism being robust enough to find those globally distributed tokens reliably under real-world load, because if that search becomes slow, all these kernel optimizations won't matter much.
Tom: So they’re essentially balancing aggressive memory reduction targets with a more complex reference search process to ensure the system actually functions well when it’s running under heavy load.
Jane: It shows they are being very thoughtful about the trade-offs; it's not just about getting the smallest number, but getting a usable number that maintains quality during generation.
Lu: The paper also highlights a need for memory-aware caching strategies within Sparse-vLLM, where they suggest managing GPU memory at the token level rather than just handling it in large pages, which is a key architectural refinement.
Meng: That token-level management sounds very promising because it means we can keep the most important compressed residuals and references right in fast GPU memory while offloading less critical data elsewhere, which reduces I/O bottlenecks.
Tom: So they’re suggesting a tiered approach: use this residual compression for efficiency, integrate it with sparse attention for selectivity, and manage the resulting memory structure intelligently at the hardware level.
Jane: It’s a really layered approach that addresses the problem from data representation all the way down to hardware execution.
Lu: And looking ahead, they suggest exploring full-pipeline quantization as a next step beyond just residual compression, which could potentially bring those memory savings much closer to theoretical limits.
Meng: I'm watching that closely; if we can get close to that seven point two percent figure with better runtime management, it would drastically improve the energy efficiency of serving these large models.
Tom: It’s a lot of moving parts, but the direction they are pointing toward—making these caches smarter and more integrated with modern inference engines—is definitely where the industry is heading.
Conclusion: Tom: So we’ve wrapped up our deep dive into ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference, which basically shows how using residual encoding based on long-range similarity can cut KV cache memory down to about twenty-nine percent while boosting throughput when paired with Sparse-vLLM.
Jane: That’s a massive win because it tackles the linear growth problem in LLMs head-on, giving us a much clearer path toward deploying powerful models in demanding, real-world applications without hitting those frustrating memory walls.
Lu: The paper's main implication is shifting our focus from simple eviction policies to similarity-aware representations for context management, which opens up whole new avenues for how we architect long-context inference systems.
Meng: From my side, the practical impact is huge; if we can deploy these models on less powerful hardware while maintaining high performance metrics like compute ratio, it drastically lowers the barrier to entry for serving state-of-the-art models at scale.
Lalam: For our internal LLM vision, this means we can support significantly longer reasoning chains and document processing tasks natively because the memory constraint that usually kills context length has been substantially loosened, which will improve the overall quality of our complex analysis capabilities.
Tom: Exactly! It’s not just about fitting things in; it’s about doing complex reasoning with high-fidelity context for much longer periods.
Jane: And they tied this compression into an inference engine called Sparse-vLLM, which lets the system selectively decompress only what’s needed for a specific attention step, which really makes it practical.
Lu: The integration with Sparse-vLLM is what elevates this beyond just a theoretical compression method; it shows how the mathematical concept of residuals can actually be operationalized in a high-throughput inference environment.
Meng: I’m interested in the performance numbers they showed, specifically that they saw up to two times higher throughput when using this combination with Sparse-vLLM compared to standard vLLM setups in long context scenarios.
Tom: That's the real kicker for deployment; a memory saving is one thing, but doubling the speed makes it usable in demanding applications like autonomous agents, which is what we’re targeting.
Jane: It really moves us toward a future where powerful LLMs aren't just experimental tools but robust infrastructure that can handle truly massive workloads efficiently.
Lu: If we think about the broader AI landscape, this suggests that focusing on similarity-aware representations across long distances is the way forward for scaling context in these models.
Meng: It forces us to reconsider how we manage memory at the token level instead of just handling it in large pages, which is a much more granular and efficient approach for serving models.
Lalam: For our culture, this means we can build systems that handle massive amounts of context reliably and efficiently, opening up new possibilities for high-fidelity document analysis across our entire platform.
Tom: So we’ve seen how DeltaKV uses residual encoding based on global similarity to achieve a twenty-nine percent memory reduction and a potential two times throughput boost when paired with Sparse-vLLM.
Jane: It really shows that by looking at the global structure of the cache, we can find intelligent ways to compress it while preserving what makes those long contexts useful for complex tasks.
Lu: The paper’s main contribution is showing a principled way to exploit this global redundancy rather than relying on local heuristics, which is a major conceptual leap in KV optimization.
Meng: We need to keep an eye on how they handle that training objective combining reconstruction loss and next-token prediction; that balance between accuracy and generation quality is crucial for real-world deployment.
Tom: And we'll be keeping a close watch on those future work directions, especially the goal of reaching even smaller footprints with full-pipeline quantization down to seven point two percent.
Jitai Hao, Qiang Huang, Yaowei Wang, Min Zhang, Jun Yu
cs.CL, cs.AI
Submitted: 2026-02-08
Updated: 2026-09-30
Code: https://github.com/CURRENTF/Sparse-vLLM
Importance score: 92/100
The gist: DeltaKV introduces a residual-based KV cache compression framework that leverages long-range inter-token similarity to significantly reduce memory footprint while maintaining near-lossless
Key concepts
- Long-Range Inter-Token Similarity
- This concept identifies that semantically similar tokens are not just close together but can be found far apart in a long context. The paper shows that over 60% of similar tokens occur beyond 16 positions, meaning compression must look globally for relevant data rather than just looking at neighbors.
- Highly Shared Latent Components
- KV caches contain common linguistic and structural patterns shared by many tokens. DeltaKV leverages this by subtracting these shared components from individual token representations. This leaves only 'residual' information, which is much smaller and easier to compress without losing critical meaning.
- Residual Computation
- The core of the compression involves calculating a residual vector (z∆) by subtracting a projected mean reference KV state from the current token's projection. This subtraction effectively isolates the unique, non-shared information for each new token, allowing only this small residual to be encoded and stored efficiently.
Terminology
Summary
DeltaKV introduces a residual-based KV cache compression framework that leverages long-range inter-token similarity to significantly reduce memory footprint while maintaining near-lossless performance for long-context Large Language Models. This work matters because it addresses the critical bottleneck of linear KV cache growth, which severely limits the deployment of powerful LLMs in demanding applications like autonomous agents and complex reasoning tasks. By identifying global redundancy in KV representations, DeltaKV provides a principled path toward scalable long-context LLM deployment by reducing cache memory to 29% of its original size and achieving up to 2x throughput improvement when integrated with Sparse-vLLM.
Key Empirical Observations
The framework is motivated by two key empirical findings concerning KV caches:
-
Long-Range Inter-Token Similarity: The paper shows that
semantically similar tokens are often distributed globally across the context rather than confined to nearby positions.
Specifically,over 60% of the most similar tokens occur at distances greater than 16,
indicating that compression must leverage global retrieval rather than local heuristics. -
Highly Shared Latent Components: KV caches exhibit strong anisotropy, with a
small number of high-norm latent directions capturing common linguistic and structural patterns shared across many tokens.
Subtracting retrieved references effectively removes these shared components, causingresidual KVs to collapse to low-magnitude, noise-like signals.
DeltaKV Framework Architecture
DeltaKV partitions the KV cache into a small set of uncompressed reference tokens and a larger set of compressed tokens. The core mechanism involves encoding semantic residuals relative to retrieved historical references. The process is formalized through the following steps:
-
Strided Reference Selection: A fixed interval stride 's' is used to maintain a reference set T, and the top-k nearest tokens from this set are retrieved based on L2 distance to compute the mean reference representation, KV R.
-
Residual Computation: The residual vector z∆ is computed by subtracting the projected mean reference from the current token's projection:
z∆ = fc(KV) − fc(KV R).
-
Reconstruction: The compressed residual codes are decoded through a decompressor fd to yield the reconstructed KV state, and finally,
The final KV cache is recovered as: KVd i = KVd ∆ + KV R.
Training and Inference Integration
DeltaKV is trained using a hybrid objective function to ensure both accuracy and generative capability. The total loss L combines a reconstruction loss (Lrec) with a next-token prediction loss (Lntp): L = Lrec + Lntp(θ, ϕ).
This ensures that the compression does not discard features essential for end-to-end generation. During inference, DeltaKV is designed to complement sparse attention methods like OmniKV. It allows for selective decompression: we only reconstruct the KV pairs required by the sparse attention mask,
avoiding unnecessary decompression and memory I/O when combined with Sparse-vLLM.
Practical Deployment via Sparse-vLLM
To enable practical deployment, the paper introduces Sparse-vLLM, an inference engine optimized for sparse and compressed KV layouts. This framework decouples memory management from execution by introducing a modular CacheManager and a centralized Sparse Controller. The controller orchestrates View Construction
(pre-forward) and Lifecycle Management
(post-forward). At the operator level, Sparse-vLLM utilizes custom Triton kernels optimized for DeltaKV, including a "Batch L2 Distance kernel for rapid reference searching and a Fused Reconstruction kernel that combines the gathering of reference tokens, mean calculation, and residual addition into a single kernel launch to minimize GPU memory bandwidth consumption."
Performance Results
Experiments on LongBench, SCBench, and AIME demonstrate that DeltaKV achieves competitive performance while being more memory-efficient than dynamic selection methods. For instance, on Llama-3.1-8B (30% budget), DeltaKV achieves a KV Cache Keep Ratio (KR) of approximately 45.2%, significantly lower than the 100% KR of OmniKV/Quest, while maintaining a comparable Compute Ratio (CR). Furthermore, when integrated with Sparse-vLLM, DeltaKV yields up to 2× higher throughput over vLLM in longcontext scenarios.
The framework is also shown to be quantization-friendly,
as the compressed residuals have a distribution that supports token-wise quantization.
Impact and Future Directions
The work concludes that KV optimization should move beyond eviction and low-rank projection toward similarity-aware representations. DeltaKV's design reduces KV cache memory to 29% while preserving near-lossless accuracy.
Future optimizations focus on reducing runtime overhead by moving view/slot management onto GPU kernels and exploring full-pipeline quantization, which could theoretically reduce the footprint to approximately 7.2% of the original size. The paper emphasizes that DeltaKV's success stems from exploiting global redundancy rather than local similarity assumptions.
Improvements for AI systems
Based on the scientific paper DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity,
here are specific, actionable improvements for AI systems and what those improved systems can achieve:
)Improved System Capabilities and Specific Enhancements:
-
[] Implement a novel, residual-based KV cache compression module that encodes token-specific deviations relative to a small set of globally retrieved historical reference tokens.
-
[] Integrate this DeltaKV framework into a high-performance inference engine (Sparse-vLLM) featuring decoupled memory management and kernels optimized for sparse and irregular KV layouts.
-
[] Develop an inference pipeline that leverages the
Sparse Controller
to dynamically construct logical views based on reference tokens, allowing selective decompression of only the necessary tokens during attention computation. -
[] Design a training objective combining Mean Squared Error (MSE) reconstruction loss with Next-Token Prediction (NTP) loss to ensure both fidelity and generative capability are preserved during compression training.
-
[] Implement an asymmetric compressor/decompressor design (SwiGLU block for compression, bias-free linear projection for decompression) to minimize runtime overhead during repeated decoding steps.
-
[] Introduce a mechanism for full-pipeline quantization, applying 4-bit quantization uniformly across the compressed KV residuals and potentially the uncompressed components in sparse layers to achieve extreme memory reduction (theoretical footprint down to 7.2% of original size).
-
[] Optimize kernel execution by utilizing indirect addressing via slot mapping and implementing fused DeltaKV kernels (combining reference gathering, mean calculation, residual application, and reconstruction) within Triton kernels to minimize global-memory traffic.
-
[] Develop a memory-aware caching strategy within the Sparse-vLLM framework that manages GPU memory at the token level rather than the page level, keeping only high-utility compressed residuals and references in fast GPU memory while offloading less critical data to host memory.
)What These Improved Systems Can Do:
-
[] Deploy extremely long context LLMs (e.g., 1M+ tokens) on consumer-grade hardware (like an RTX PRO 6000) without running out of GPU memory, enabling true
infinite
context capabilities with minimal latency penalties. -
[] Achieve up to a 2× throughput improvement over existing inference engines (like vLLM) in long-context scenarios, significantly increasing the speed of response generation for complex reasoning and document analysis tasks.
-
[] Enable real-time, high-fidelity execution of complex multi-turn dialogue systems (SCBench tasks) and intricate mathematical reasoning benchmarks (AIME), which currently suffer from performance degradation when using static eviction methods.
-
[] Facilitate the practical deployment of large, state-of-the-art models in resource-constrained environments by drastically reducing the memory footprint (down to 29% original size) while maintaining near-lossless accuracy across critical tasks like Code Generation and Retrieval KV (R.KV).
-
[] Significantly reduce energy consumption and PCIe bandwidth bottlenecks associated with moving large KV cache data between GPU and host memory during inference, making large-scale LLM serving more energy-efficient.
Sources
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- Training Deep Nets with Sublinear Memory Cost
- Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Enhancing RAG Efficiency with Adaptive Context Compression
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- CentroidKV: Efficient Long-Context LLM Inference via KV Cache Clustering
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
- SCBench: A KV Cache-Centric Analysis of Long-Context Methods
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
- Transformers are Multi-State RNNs
- DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference
- Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering