G squared PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
cs.CL, cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/G2PTQ/G2PTQ
Terminology
Sources
- DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
- The Llama 3 Herd of Models
- Statistically-Lossless Quantization of Large Language Models
- OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
- Identifying Sensitive Weights via Post-quantization Integral
- C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
- GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
- BoA: Attention-aware Post-training Quantization without Backpropagation
- TurboBoA: Faster and Exact Attention-aware Quantization without Backpropagation
- GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration
- IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact
- SpinQuant: LLM quantization with learned rotations
- What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
- Pointer Sentinel Mixture Models
- A White Paper on Neural Network Quantization
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering