VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits
cs.CL
Submitted: 2025-05-15
Updated: 2026-09-18
Comments: Lack of sufficient experiments and detailed format alignment
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Language Models are Few-Shot Learners
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- Learning both Weights and Connections for Efficient Neural Networks
- Distilling the Knowledge in a Neural Network
- Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
- Categorical Reparameterization with Gumbel-Softmax
- On Using Very Large Target Vocabulary for Neural Machine Translation
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Decoupled Weight Decay Regularization
- Using the Output Embedding to Improve Language Models
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Neural Discrete Representation Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering