GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression
cs.LG
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/tatsu-lab/stanford_alpaca
License: http://creativecommons.org/licenses/by/4.0/
The gist: Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard
Terminology
Abstract
Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations. Rather than forcing weights of adjacent layers to share a basis or heuristically merging activation statistics, our approach identifies structurally compatible projections and learns a shared representation that better preserves each layer's distinct calibration geometry. Coupled with structured sparsity, this yields highly efficient weight decompositions without sacrificing functional fidelity. Across diverse architectures, scales, and modalities, our method achieves state-of-the-art results, consistently outperforming independent structured weight decompositions and alternative pairwise weight factorizations, which operate under heuristic grouping strategies. By replacing heuristic engineering strategies with a convergent, optimization-driven pipeline, we establish a theoretically grounded foundation for scalable, transformer compression across different modalities.
Sources
- Phi-4 Technical Report
- ROCKET: Rapid Optimization via Calibration-guided Knapsack Enhanced Truncation for Efficient Model Compression
- CoSpaDi: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- COMPOT: Calibration-Optimized Matrix Procrustes Orthogonalization for Transformers Compression
- gpt-oss-120b & gpt-oss-20b Model Card
- VibeVoice Technical Report
- MLS: A Large-Scale Multilingual Dataset for Speech Research
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Gemma 3 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- Qwen3 Technical Report
- Instruction-Following Evaluation for Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks