Tensor Decomposition of Transformer Key-Value Caches: Spectral Structure and Format Comparison
math.NA, cs.CL, cs.LG, cs.NA
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/rahulk98/JoLT-Master-Thesis
Terminology
Sources
- Fast Transformer Decoding: One Write-Head is All You Need
- Mistral 7B
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- A JoLT for the KV cache: Near-Lossless KV Cache Compression via Joint Rank-bit Allocation
Related papers
- Do physics-informed neural networks (PINNs) need to be deep? Shallow PINNs using the Levenberg-Marquardt algorithm
- A Neural-preconditioned Poisson Solver for Mixed Dirichlet and Neumann Boundary Conditions
- Second-order consistency for learning chaotic dynamics via randomized Jacobian matching
- Windowed thinning and query complexity for the bouncy particle and Zigzag samplers
- Data-efficient Kernel Methods for Learning Hamiltonian Systems
- Adjoint Method versus Physics-Informed Neural Networks in PDE-Constrained Inverse Problems