Low-Rank Ternary Adaptation for Fine-Tuning Transformers
cs.CV, cs.LG
Submitted: 2026-08-25
Updated: 2026-08-25
Comments: Accepted at ECCV 2026. To be published in Volume 17015 of the Lecture Notes in Computer Science series
Code: https://github.com/alexmanoo/ternary_adaptation
License: http://creativecommons.org/licenses/by/4.0/
The gist: Ternary transformers offer extreme memory and compute efficiency, but existing low-bit LoRA-based methods cannot directly fine-tune ternary weights.
Terminology
Abstract
Ternary transformers offer extreme memory and compute efficiency, but existing low-bit LoRA-based methods cannot directly fine-tune ternary weights. Current approaches either require dequantization, restoring low-bit base weights to higher precision to merge with adaptation weight, or update only quantization parameters, preventing a merged model that remains ternary. We propose ternary multiplicative adaptation, which represents discrete updates of ternary weights such as sign flips or zeroing through a low-rank Kronecker factorization into two small ternary matrices applied element-wise to ternary weights. This design is parameter-efficient and expressive, preserves the ternary domain, and supports direct merging without dequantization. Experiments on six models across language and vision, including ternarized LLaMA-3 1B and 3B and a ternary ViT-B/16, demonstrate that our method recovers much of the performance lost to quantization and outperforms strong low-bit and ternary baselines. Code is available at https://github.com/alexmanoo/ternary adaptation.
Sources
- LoRMA: Low-Rank Multiplicative Adaptation for LLMs
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
- BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
- KronA: Parameter Efficient Tuning with Kronecker Adapter
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Measuring Massive Multitask Language Understanding
- The Power of Scale for Parameter-Efficient Prompt Tuning
- Ternary Weight Networks
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models
- ReLoRA: High-Rank Training Through Low-Rank Updates
- ApiQ: Finetuning of 2-Bit Quantized Large Language Model
- LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
- SpinQuant: LLM quantization with learned rotations
- BitNet b1.58 2B4T Technical Report
- The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models