HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression
Yuefeng Zhang
Beijing Institute of Computer Technology and Application · Northwestern Polytechnical University
cs.CV, cs.AI, cs.MM
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: Learned image compression, post-training quantization, mixed-precision quantization, Hessian-based sensitivity analysis, model compression
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 75/100
The gist: HAMP-LIC is a Hessian-aware mixed-precision post-training quantization (PTQ) framework proposed for learned image compression (LIC) models.
Terminology
Summary
HAMP-LIC is a Hessian-aware mixed-precision post-training quantization (PTQ) framework proposed for learned image compression (LIC) models. The paper states: To enable efficient and accurate low-bit deployment of pre-trained LIC models, we propose HAMP-LIC, a Hessian-aware mixed-precision post-training quantization (PTQ) framework with a four-stage optimization strategy.
The framework consists of four stages: "First, block-wise sensitivity is estimated from the Hessian trace to capture second-order importance. Second, a task-aware refinement module adjusts these sensitivities by jointly considering quantization distortion and rate–distortion performance. Third, guided by the refined sensitivity profile, bit-widths are allocated under a global model-size constraint to balance efficiency and reconstruction quality. Finally, block-wise reconstruction on a small calibration set further suppresses quantization error."
The paper motivates the work by noting that Learned image compression (LIC) models achieve strong rate–distortion performance but are hindered by high computational complexity and encoding–decoding mismatches across heterogeneous hardware platforms.
It also notes that Uniform fixed-precision quantization alleviates these issues but suffers severe quality degradation at low bit-widths, as it ignores the differing quantization sensitivities of individual layers.
The main contributions are summarized as: "1) We propose HAMP-LIC, a mixed-precision PTQ framework that integrates Hessian-based second-order sensitivity estimation with rate–distortion-aware optimization for LIC models. 2) We design a task-aware sensitivity refinement that modulates the Hessian-trace metric with the rate–distortion loss degradation of each block, so that bit-width decisions directly reflect the compression objective rather than generic reconstruction error. 3) We formulate bit allocation as a constrained integer optimization problem and introduce an efficient Pareto-frontier search strategy that reduces complexity from exponential to tractable scale, combined with progressive block-wise refinement. 4) We demonstrate that HAMP-LIC achieves up to 4.85× model compression with negligible BD-rate degradation, while ensuring cross-platform numerical consistency and eliminating floating-point-induced decoding mismatches."
In Step 1, the paper describes sensitivity estimation: "A straightforward choice for measuring block sensitivity is the first-order gradient information. However, first-order metrics are often insufficient for accurately capturing layer-wise sensitivity, as they do not account for the curvature of the loss landscape. To address this limitation, we adopt second-order information as the primary sensitivity metric. Specifically, we utilize the Hessian trace of the loss function with respect to the weights of each block... To ensure computational tractability, we instead estimate the Hessian trace efficiently via Hessian-vector products combined with randomized probing techniques, based on Hutchinson's method."
In Step 2, the task-aware sensitivity is defined: "For each block Bl we define a task-aware sensitivity Ωl as a composite measure that integrates both the potential sensitivity captured by the Hessian trace and the task-oriented loss degradation induced by quantization. Formally, Ωl is expressed as: Ωl(bl) = Tr(Hl) · Lqp(bl) − Lint8 / Lint8, where L is the number of quantizable blocks, bl is the weight bit-width assigned to the l-th block, Tr(Hl) is the Hessian trace of block Bl computed in Step 1... The term Lqp(bl) denotes the task loss evaluated when the weights of block Bl are quantized to bl bits, while Lint8 denotes the task loss of the reference model whose weights are uniformly quantized to 8 bits."
In Step 3, the bit allocation is formulated: "Given the task-aware sensitivities Ωl refined in Step III-C, this step solves the constrained optimization problem... argmin Σ Ωl(bl) subject to (1/8L) Σ bl ≤ ϵ, where ϵ is a hyperparameter controlling the overall compression ratio relative to 8-bit quantization, which is empirically set to ϵ = 0.75 in our experiments. To reduce search complexity, the paper states:
Blocks are first sorted in descending order of their sensitivity scores, and the search is reformulated as a partition problem over the sorted list, restricting assignments to monotonically non-increasing order with respect to sensitivity rank... This formulation reduces the search space from exponential to polynomial in L, making the search computationally feasible. For the same example (L = 50, k = 3), Eq. (7) yields B = 3 + 147 + 1176 = 1326, a reduction of roughly 21 orders of magnitude compared with the exhaustive search space of 3 50 ≈ 7.12 × 10 23."
In Step 4, block-wise optimization is performed: "After acquiring the optimal bit-width allocation for each block in Step 3, to further improve the quantization performance, we adopt task-constraint block-wise optimization, which can be divided by optimization target into two parts: scaling optimization and rounding optimization. The scaling optimization minimizes
the discrepancy between the quantized and full-precision rate-distortion loss, and the rounding optimization
adopts an adaptive rounding strategy in which a learnable variable V is introduced to parameterize the rounding of weights."
Experiments were conducted on two representative LIC models... Minnen2018 and Cheng2020,
using a calibration dataset, which only consists of 12 images randomly selected from the CLIC dataset.
The results show: "Experiments on representative LIC models, including Minnen2018 and Cheng2020, demonstrate that HAMP-LIC achieves up to 4.85× model compression with as low as 0.59% BD-rate loss, consistently outperforming existing fixed- and mixed-precision PTQ methods across multiple datasets while completely eliminating cross-platform encoding–decoding errors."
Specifically, in the BD-rate comparison table, "the proposed HAMP-LIC achieves the highest model compression ratio (4.85×) among all compared methods. With full-precision activations (w=6.6, a=32), it also attains the lowest BD-rate loss on the Cheng2020 model (0.59% on Kodak and 1.79% on Tecnick), outperforming the mixed-precision FMPQ baseline (0.89% and 2.68%) despite the more aggressive weight compression."
Regarding cross-platform robustness, the paper reports: "As shown in Table II, the FP32 model suffers from significant decoding errors when encoding and decoding are performed on different platforms (CPU/GPU or GPU/CPU)... In contrast, the Proposed HAMP-LIC method reduces the error rate to zero on both datasets and for both cross-platform scenarios, ensuring reliable decoding regardless of the hardware used for encoding or decoding."
The ablation study on sensitivity metrics shows: "it reveals that utilizing the Hessian matrix yields superior results compared to the Fisher matrix under the same bit-rate restriction. Second-order metrics like the Hessian are superior to first-order gradients because they capture the curvature of the loss landscape, not just its slope."
The bit allocation ablation shows: "the proposed method achieves a lower average bit-width of 5.719 with ϵ set to 0.65. Experimental results in Figure 5 show that the proposed method maintains rate distortion performance consistent with the baseline, even outperforming it at higher bitrates (bpp > 0.6)."
The bit distribution analysis reveals: "We observe that main-path weights (layers [0–40]) are generally more sensitive to quantization, requiring higher bit-widths than hyper-path weights. Specifically, the last layers of the encoder (layer 18) and decoder (layer 39) necessitate high bit-widths, as they encode critical information regarding the latent representation and the reconstructed output image, respectively."
The BOPS analysis shows: "our mixed-precision quantization achieves BOPS of 2007.5 G, 4628.2 G, 4652.6 G, and 4661.9 G for quality levels λ ∈ 0.0067, 0.025, 0.025, 0.0483, corresponding to reductions of 94.3%, 94.1%, 94.1%, and 94.1% relative to the full-precision (FP32) baseline. Compared to uniform 8-bit weight quantization (W8A8), our method further reduces BOPS by 8.6%, 6.0%, 5.5%, and 5.3%, respectively."
The paper concludes: "In this paper, we presented HAMP-LIC, a Hessian-aware mixed-precision post-training quantization framework for deploying learned image compression models on resource-constrained hardware. HAMP-LIC combines Hessian-trace-based block sensitivity with a task-aware rate–distortion criterion, allocates bit widths through an efficient Pareto-frontier search, and applies block-wise reconstruction to reduce quantization error. Experiments on the Minnen2018 and Cheng2020 models across multiple benchmark datasets show that HAMP-LIC achieves up to 4.85× model compression with negligible BD-rate loss, outperforms existing fixed- and mixed-precision PTQ methods, and completely eliminates cross-platform encoding–decoding errors. Future work will extend HAMP-LIC to transformer-based LIC architectures, automate the selection of the compression-ratio hyperparameter ϵ, and reduce calibration overhead."
Improvements for AI systems
Improvements to AI Systems:
-
Hessian-Aware Sensitivity Estimation for Quantization: Integrate Hessian-trace-based second-order sensitivity (via Hutchinson’s method) into any neural network’s post-training quantization pipeline, replacing first-order gradient metrics. This enables more accurate identification of layers that are critical to task performance, reducing quality degradation at low bit-widths.
-
Task-Aware Sensitivity Refinement: Modify the sensitivity metric to combine Hessian trace with task-specific loss degradation (e.g., rate–distortion loss for compression, or classification loss for vision models). This ensures bit-width allocation directly optimizes the end-task objective rather than generic reconstruction error, improving performance on domain-specific metrics.
-
Efficient Pareto-Frontier Bit Allocation: Implement the constrained integer optimization with monotonic bit-width assignment over sensitivity-sorted blocks. This reduces search complexity from exponential to polynomial, allowing real-time bit allocation for large models (e.g., 50+ blocks) on edge devices without exhaustive search.
-
Progressive Block-Wise Reconstruction with Adaptive Rounding: Apply block-wise scaling and learnable rounding optimization on a small calibration set (e.g., 12 images). This suppresses quantization error without full retraining, enabling rapid deployment of compressed models on new hardware.
-
Cross-Platform Numerical Consistency: Enforce deterministic integer-only operations (as implied by the framework) to eliminate floating-point-induced encoding–decoding mismatches. This guarantees identical outputs across CPU/GPU and heterogeneous hardware, critical for distributed or edge-cloud pipelines.
-
Compression-Ratio-Aware Deployment: Use the hyperparameter ϵ to control model size vs. quality trade-off dynamically. The system can automatically adjust bit-widths to meet a target model-size constraint (e.g., 4.85× compression) while maintaining BD-rate loss below 1%, enabling adaptive deployment on memory-limited devices.
What the Improved AI System Can Do:
-
Compress pre-trained image compression models (e.g., Minnen2018, Cheng2020) by up to 4.85× with <0.6% BD-rate loss, outperforming uniform 8-bit and existing mixed-precision methods.
-
Deploy on heterogeneous hardware (CPU/GPU/edge) with zero cross-platform decoding errors, ensuring reliable inference in production.
-
Allocate bit-widths in under a second for models with 50 layers, making it suitable for real-time model optimization on-device.
-
Maintain rate–distortion performance even at aggressive compression (average bit-width 5.7 bits), sometimes outperforming the full-precision baseline at higher bitrates.
-
Reduce computational cost (BOPS) by 94% relative to FP32 and an additional 5–8% over uniform 8-bit quantization, enabling real-time inference on resource-constrained devices.
-
Extend the same framework to other tasks (e.g., classification, detection) by substituting the task loss, providing a general-purpose Hessian-aware PTQ tool.
Abstract
Use this plain-text version for the arXiv abstract field: Learned image compression (LIC) models achieve strong rate-distortion performance but are hindered by high computational complexity and encoding-decoding mismatches across heterogeneous hardware platforms. Uniform fixed-precision quantization alleviates these issues but suffers severe quality degradation at low bit widths because it ignores differences in the quantization sensitivities of individual layers. To enable efficient and accurate low-bit deployment of pretrained LIC models, we propose HAMP-LIC, a Hessian-aware mixed-precision post-training quantization (PTQ) framework with a four-stage optimization strategy. First, block-wise sensitivity is estimated from the Hessian trace to capture second-order importance. Second, a task-aware refinement module adjusts these sensitivities by jointly considering quantization distortion and rate-distortion performance. Third, guided by the refined sensitivity profile, bit widths are allocated under a global model-size constraint to balance efficiency and reconstruction quality. Finally, block-wise reconstruction using a small calibration set further suppresses quantization error. Experiments on representative LIC models, including Minnen2018 and Cheng2020, demonstrate that HAMP-LIC achieves up to 4.85x model compression with as little as 0.59% BD-rate loss. It consistently outperforms existing fixed- and mixed-precision PTQ methods across multiple datasets while completely eliminating cross-platform encoding-decoding errors.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models